Model reference · open weights
IQuest-Q1 is an open-weight language model from IQuestLab. IQuest-Q1 (BF16) weighs 641 GB; the smallest configuration that runs it is 4× B200 180 GB.
What it is
| Released by | IQuestLab |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 320.3B |
| Context | 524,288 tokens |
| Runs with | transformers |
| Released | 2026-09-28 |
| Popularity | 503 downloads / month |
| Weights | 641 GB (IQuest-Q1 (BF16), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 641 GB (file size) · KV cache 360 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 656 MB on a small card · context up to 524,288 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … 8× A100 80 GB 14 smaller cards | — | — | — | |
| 4× B200 180 GB tensor parallel | 14 | 3 | 119K | 176 GB a card |
| 8× H200 141 GB tensor parallel | 143 | 35 | all 512K | 138 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 644 GB | 653 GB |
| 5 | 656 GB | 700 GB |
| 8 | 665 GB | 736 GB |
| 16 | 689 GB | 830 GB |
| 32 | 736 GB | 1019 GB |
| 64 | 830 GB | 1397 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
IQuest-Q1 is a Mixture-of-Experts (MoE) model developed by IQuest for agentic coding, reasoning, and multi-step tool use. It comprises approximately 320B total parameters, with an estimated 15B parameters activated per token.
| Property | IQuest-Q1 |
|---|---|
| Total Parameters | 320B |
| Activated Parameters | 15B |
| Transformer layers | 88 |
| Hidden dimension | 3,072 |
| Attention Heads (Q/KV) | 48/8 |
| Head Dimension | 128 |
| Hybrid Attention Pattern | 3 SWA + 1 FA |
| Sliding Window Size | 4,096 |
| Partial RoPE Dimensions | 32 |
| Experts (Total/Activated) | 256/8 |
| MTP Layers | 2 Independent (Training) / 1 Recursive x8 (Inference) |
| MTP Sliding Window Size | 512 (Inference) |
| Context length | 524,288 |
For reproducibility, we recommend using a temperature of 1.0, top-p of 0.95, and top-k of 20, with Claude Code 2.1.140 or Codex 0.142 as the respective harness.
For each model, we report the publicly reported score; otherwise, we evaluate the model using the corresponding benchmark setup:
(1) Harness: For agentic coding tasks, we use mini-SWE-agent for DeepSWE v1.1, and Claude Code for other tasks. For Agents' Last Exam, we evaluate our model using Claude Code 2.1.258. Because our models currently do not have multimodal capability, we replace multimodal content inputs in the agent conversation with placeholders during tokenization. (2) Runtime: We set six-hour limits for CyberGym, eight hours for Terminal-Bench 2.1.
Start a server as described in Deployment, install the client with pip install openai, and then call the OpenAI-compatible API:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="sk-iquest")
response = client.chat.completions.create(
model="IQuest-Q1",
messages=[
{"role": "user", "content": "Hello! Can you briefly introduce yourself?"},
],
temperature=1.0,
top_p=0.95,
)
print(response.choices[0].message.content)
For production deployment, we recommend using SGLang or vLLM.
Use our prebuilt image iquestlabworkspace/sglang-iquest-q1:cu130:
docker pull iquestlabworkspace/sglang-iquest-q1:cu130
# without MTP
MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
python -u -m sglang.launch_server \
--model-path "$MODEL_ROOT" \
--served-model-name IQuest-Q1 \
--tp-size 8 \
--dtype bfloat16 \
--attention-backend fa3 \
--mem-fraction-static 0.85 \
--disable-prefill-cuda-graph \
--enable-metrics \
--tool-call-parser iquest_q1 \
--reasoning-parser iquest_q1 \
--enable-torch-compile \
--load-format fastsafetensors \
--speculative-use-rejection-sampling
# with recursive MTP
MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
python -u -m sglang.launch_server \
--model-path "$MODEL_ROOT" \
--served-model-name IQuest-Q1 \
--tp-size 8 \
--dtype bfloat16 \
--attention-backend fa3 \
--mem-fraction-static 0.85 \
--disable-prefill-cuda-graph \
--enable-metrics \
--tool-call-parser iquest_q1 \
--reasoning-parser iquest_q1 \
--speculative-algorithm EAGLE \
--speculative-num-steps 5 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 6 \
--speculative-draft-model-path "$MODEL_ROOT/mtp" \
--enable-torch-compile \
--load-format fastsafetensors \
--speculative-use-rejection-sampling \
--speculative-draft-attention-backend fa3
Use our image iquestlabworkspace/vllm-iquest-q1:cu130:
docker pull iquestlabworkspace/vllm-iquest-q1:cu130
# without MTP
MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
vllm serve "$MODEL_ROOT" \
--served-model-name IQuest-Q1 \
--tensor-parallel-size 8 \
--reasoning-parser iquest_q1 \
--enable-auto-tool-choice \
--tool-call-parser iquest_q1
# with recursive MTP
MODEL_ROOT="$(hf download IQuestLab/IQuest-Q1 --quiet)" && \
vllm serve "$MODEL_ROOT" \
--served-model-name IQuest-Q1 \
--tensor-parallel-size 8 \
--reasoning-parser iquest_q1 \
--enable-auto-tool-choice \
--tool-call-parser iquest_q1 \
--enable-prefix-caching \
--speculative-config '{
"method": "eagle",
"model": "'"$MODEL_ROOT"'/mtp",
"num_speculative_tokens": 5,
"draft_sample_method": "probabilistic",
"rejection_sample_method": "standard",
"enforce_eager": false
}'
Use a gateway that supports tool calling via Anthropic Messages for Claude Code and OpenAI Responses for Codex.
We recommend version 2.1.140 and the IQuest-Q1[1m] client-side setting does not change 512K context limit.
export ANTHROPIC_MODEL="IQuest-Q1[1m]"
export ANTHROPIC_DEFAULT_SONNET_MODEL="IQuest-Q1[1m]"
export ANTHROPIC_DEFAULT_OPUS_MODEL="IQuest-Q1[1m]"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="IQuest-Q1[1m]"
export CLAUDE_CODE_SUBAGENT_MODEL="IQuest-Q1[1m]"
export CLAUDE_CODE_MAX_OUTPUT_TOKENS="131072"
export CLAUDE_AUTOCOMPACT_PCT_OVERRIDE="80"
export CLAUDE_CODE_AUTO_COMPACT_WINDOW="524288"
export API_TIMEOUT_MS="3000000"
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC="1"
export CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS="1"
export ANTHROPIC_BASE_URL="http://example-iquest-q1-link"
export ANTHROPIC_AUTH_TOKEN="sk-iquest-q1"
claude --model IQuest-Q1
We recommend version 0.142.0.
export OPENAI_API_KEY="sk-iquest-q1"
export MODEL_ID="IQuest-Q1"
# Add /v1 if required by your gateway's Responses endpoint.
export BASE_URL="http://example-iquest-q1-link"
(
set -eu
CONFIG_DIR="${HOMQuoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.