Model reference · open weights

IQuest-Q1

NEW · this week LLMs IQuestLab Text gen 1 build Its own licence terms 503 dl/mo

IQuest-Q1 is an open-weight language model from IQuestLab. IQuest-Q1 (BF16) weighs 641 GB; the smallest configuration that runs it is 4× B200 180 GB.

What it is

Released byIQuestLab
TypeLanguage models
TaskText gen
Parameters (lead)320.3B
Context524,288 tokens
Runs withtransformers
Released2026-09-28
Popularity503 downloads / month
Weights641 GB (IQuest-Q1 (BF16), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for IQuest-Q1 (BF16)

Weights 641 GB (file size) · KV cache 360 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 656 MB on a small card · context up to 524,288 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … 8× A100 80 GB
14 smaller cards
———
4× B200 180 GB
tensor parallel
143119K176 GB a card
8× H200 141 GB
tensor parallel
14335all 512K138 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
1644 GB653 GB
5656 GB700 GB
8665 GB736 GB
16689 GB830 GB
32736 GB1019 GB
64830 GB1397 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What IQuestLab says about IQuest-Q1


Introduction

IQuest-Q1 is a Mixture-of-Experts (MoE) model developed by IQuest for agentic coding, reasoning, and multi-step tool use. It comprises approximately 320B total parameters, with an estimated 15B parameters activated per token.

Read the full model card

Model Specifications

PropertyIQuest-Q1
Total Parameters320B
Activated Parameters15B
Transformer layers88
Hidden dimension3,072
Attention Heads (Q/KV)48/8
Head Dimension128
Hybrid Attention Pattern3 SWA + 1 FA
Sliding Window Size4,096
Partial RoPE Dimensions32
Experts (Total/Activated)256/8
MTP Layers2 Independent (Training) / 1 Recursive x8 (Inference)
MTP Sliding Window Size512 (Inference)
Context length524,288

Performance

Benchmark Notes

For reproducibility, we recommend using a temperature of 1.0, top-p of 0.95, and top-k of 20, with Claude Code 2.1.140 or Codex 0.142 as the respective harness.

For each model, we report the publicly reported score; otherwise, we evaluate the model using the corresponding benchmark setup: (1) Harness: For agentic coding tasks, we use mini-SWE-agent for DeepSWE v1.1, and Claude Code for other tasks. For Agents' Last Exam, we evaluate our model using Claude Code 2.1.258. Because our models currently do not have multimodal capability, we replace multimodal content inputs in the agent conversation with placeholders during tokenization. (2) Runtime: We set six-hour limits for CyberGym, eight hours for Terminal-Bench 2.1.

Quick Start

Start a server as described in Deployment, install the client with pip install openai, and then call the OpenAI-compatible API:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="sk-iquest")
response = client.chat.completions.create(
    model="IQuest-Q1",
    messages=[
        {"role": "user", "content": "Hello! Can you briefly introduce yourself?"},
    ],
    temperature=1.0,
    top_p=0.95,
)
print(response.choices[0].message.content)

Deployment

For production deployment, we recommend using SGLang or vLLM.

SGLang

Use our prebuilt image iquestlabworkspace/sglang-iquest-q1:cu130:

docker pull iquestlabworkspace/sglang-iquest-q1:cu130

# without MTP
MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
python -u -m sglang.launch_server \
  --model-path "$MODEL_ROOT" \
  --served-model-name IQuest-Q1 \
  --tp-size 8 \
  --dtype bfloat16 \
  --attention-backend fa3 \
  --mem-fraction-static 0.85 \
  --disable-prefill-cuda-graph \
  --enable-metrics \
  --tool-call-parser iquest_q1 \
  --reasoning-parser iquest_q1 \
  --enable-torch-compile \
  --load-format fastsafetensors \
  --speculative-use-rejection-sampling

# with recursive MTP
MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
python -u -m sglang.launch_server \
    --model-path "$MODEL_ROOT" \
    --served-model-name IQuest-Q1 \
    --tp-size 8 \
    --dtype bfloat16 \
    --attention-backend fa3 \
    --mem-fraction-static 0.85 \
    --disable-prefill-cuda-graph \
    --enable-metrics \
    --tool-call-parser iquest_q1 \
    --reasoning-parser iquest_q1 \
    --speculative-algorithm EAGLE \
    --speculative-num-steps 5 \
    --speculative-eagle-topk 1 \
    --speculative-num-draft-tokens 6 \
    --speculative-draft-model-path "$MODEL_ROOT/mtp" \
    --enable-torch-compile \
    --load-format fastsafetensors \
    --speculative-use-rejection-sampling \
    --speculative-draft-attention-backend fa3

vLLM

Use our image iquestlabworkspace/vllm-iquest-q1:cu130:

docker pull iquestlabworkspace/vllm-iquest-q1:cu130

# without MTP
MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
vllm serve "$MODEL_ROOT" \
  --served-model-name IQuest-Q1 \
  --tensor-parallel-size 8 \
  --reasoning-parser iquest_q1 \
  --enable-auto-tool-choice \
  --tool-call-parser iquest_q1

# with recursive MTP
MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
vllm serve "$MODEL_ROOT" \
  --served-model-name IQuest-Q1 \
  --tensor-parallel-size 8 \
  --reasoning-parser iquest_q1 \
  --enable-auto-tool-choice \
  --tool-call-parser iquest_q1 \
  --enable-prefix-caching \
  --speculative-config '{
    "method": "eagle",
    "model": "'"$MODEL_ROOT"'/mtp",
    "num_speculative_tokens": 5,
    "draft_sample_method": "probabilistic",
    "rejection_sample_method": "standard",
    "enforce_eager": false
  }'

Tool Use and Agent Integration

Use a gateway that supports tool calling via Anthropic Messages for Claude Code and OpenAI Responses for Codex.

Claude Code

We recommend version 2.1.140 and the IQuest-Q1[1m] client-side setting does not change 512K context limit.

export ANTHROPIC_MODEL="IQuest-Q1[1m]"
export ANTHROPIC_DEFAULT_SONNET_MODEL="IQuest-Q1[1m]"
export ANTHROPIC_DEFAULT_OPUS_MODEL="IQuest-Q1[1m]"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="IQuest-Q1[1m]"
export CLAUDE_CODE_SUBAGENT_MODEL="IQuest-Q1[1m]"
export CLAUDE_CODE_MAX_OUTPUT_TOKENS="131072"
export CLAUDE_AUTOCOMPACT_PCT_OVERRIDE="80"
export CLAUDE_CODE_AUTO_COMPACT_WINDOW="524288"
export API_TIMEOUT_MS="3000000"
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC="1"
export CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS="1"
export ANTHROPIC_BASE_URL="http://example-iquest-q1-link"
export ANTHROPIC_AUTH_TOKEN="sk-iquest-q1"
claude --model IQuest-Q1
Codex CLI

We recommend version 0.142.0.

export OPENAI_API_KEY="sk-iquest-q1"
export MODEL_ID="IQuest-Q1"
# Add /v1 if required by your gateway's Responses endpoint.
export BASE_URL="http://example-iquest-q1-link"

(
  set -eu
  CONFIG_DIR="${HOM

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms