Model reference · open weights

AstaBrief

LLMs allenai Text gen 1 build Open weights 880 dl/mo

AstaBrief is an open-weight language model from allenai. AstaBrief_8B (BF16) weighs 16.4 GB; the smallest configuration that runs it is 2× RTX 3060 12 GB.

  • AstaBrief-8B is an 8B parameter text-generation model developed by Allen AI that converts research questions and scientific literature excerpts into cited reports.
  • It is based on Qwen3-8B and supports a context length of 40960 tokens.
  • The model is licensed under Apache 2.0 and operates in English.

Summary of the allenai/AstaBrief_8B model card, 2026-10-04

What it is

Released byallenai
Released2026-02-09
VRAM16.4 GB for the weights

What it runs on

Memory and cards for AstaBrief_8B (BF16)

16.4 GBweights, file size
147 MBcache per 1K tokens
753 MBruntime overhead, at least
40,960 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB … RTX 4060 Ti 16 GB
2 smaller cards
———
RTX 3090 24 GB51all 40K23.4 GB
RTX 4090 24 GB51all 40K23.4 GB
RTX 5090 32 GB112all 40K31.0 GB
L40S 48 GB225all 40K44.0 GB
A100 80 GB5012all 40K78.2 GB
H100 80 GB4611all 40K78.1 GB
RTX PRO 6000 Blackwell 96 GB5914all 40K93.8 GB
DGX Spark (GB10) 128 GB unified7117all 40K107 GB
H200 141 GB9624all 40K138 GB
B200 180 GB12731all 40K176 GB
2× RTX 3060 12 GB
tensor parallel
4135K11.6 GB a card
2× RTX 4060 Ti 16 GB
tensor parallel
102all 40K15.4 GB a card
2× RTX 4090 24 GB
tensor parallel
235all 40K23.4 GB a card
2× RTX 3090 24 GB
tensor parallel
235all 40K23.4 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
118.3 GB22.0 GB
523.2 GB41.3 GB
826.8 GB55.8 GB
1636.5 GB94.4 GB
3255.8 GB172 GB
6494.4 GB326 GB

One card, with vLLM's small-card settings.

From the model card

What allenai says about AstaBrief

Read the model card

AstaBrief-8B

AstaBrief-8B is a model trained to turn a research question and retrieved scientific literature excerpts into a cited report using only supervised fine-tuning and offline direct preference optimization.

The model is initialized from AstaBrief-8B-SFT and fine-tuned on AstaBrief_DPO_Mix, a collection of real user queries with two reports per query and preference judgements over those report pairs. Reports are generated from both multi-step Asta ScholarQA and single-step report generation pipelines, using a variety of backing models: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, GPT-4.1, DeepSeek-V3 and DeepSeek-R1. Two judge models – GPT-4.1 and DeepSeek-R1 – compared each pair and picked a winner. We ensured that LLM judges were aligned with human preferences (95% agreement) and only kept pairs where both judges agreed.

For a detailed overview of the project, see our blog.

Inference and Usage

[!NOTE] Recommended prompt: This checkpoint was fine-tuned using this prompt. For best results, we recommend using the same prompt format at inference time with your input query and section references. Using a different prompt or interaction format may lead to degraded or inconsistent behavior.

Following is an example inference code snippet:

from transformers import AutoModelForCausalLM, AutoTokenizer
from vllm import LLM, SamplingParams

model_name = "allenai/AstaBrief_8B"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)

# prepare the model input by providing your query and retrieved literature excerpts in our recommended prompt format
messages = [
    {"role": "user", "content": formatted_sft_prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

# instantiate sampling parameters
llm = LLM(model=model_name)
sampling_params = SamplingParams(
    temperature=0.7,
    top_p=0.95,
    max_tokens=4096,
    stop_token_ids=[tokenizer.eos_token_id],
)

# conduct text completion
outputs = llm(text, sampling_params)
output_text = outputs[0].outputs[0].text.strip()
output_ids = len(outputs[0].outputs[0].token_ids)

print("content:", output_text)

Evaluation Results

Our DPO training process improves performance over the SFT checkpoint as well as the base Qwen3-8B model on the ScholarQA-CS2 test set, a set of 100 user-written computer science research questions.

ModelAverageIngredient RecallAnswer PrecisionCitation PrecisionCitation Recall
Qwen3-8B77.377.890.676.264.6
AstaBrief-8B-SFT83.785.290.487.771.3
AstaBrief-8B8790.28990.578.2

Overall, AstaBrief-8B is competitive with Asta ScholarQA and DR-Tulu on a variety of benchmarks.

ModelSQA-CS2 (Dev)SQA-CS2 (Test)DeepScholarBenchWin Rate vs Asta SQA (SQA-CS2 Dev)Win Rate vs Asta SQA (SQA-CS2 Test)
Asta ScholarQA87.686.260.25N/AN/A
DR-Tulu-8B86.588.856.2636%54%
AstaBrief-8B86.387.053.5055%72%

Intended Uses and Limitations

This model is licensed under Apache 2.0 and is based on Qwen 3-8B. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Please refer to the details in our DPO dataset for more information about the data sources used in DPO training.

Training

DPO training was conducted in open-instruct on 8xH100s. Hyperparameter settings:

HyperparameterValue
#Epochs7
LR5e-06
LR SchedulerLinear
Warmup Ratio0.1
DPO Loss TypeDPO Norm
DPO Beta10
Per-device Batch Size1
Gradient Accumulation Steps8
Max sequence length16000
PrecisionBF16

Links

📝 Project blog: https://allenai.org/blog/astabrief

🤗 DPO dataset: AstaBrief_DPO_Mix

🤖 Base model: Qwen3-8B

🤖 SFT checkpoint: AstaBrief_8B_SFT

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms