Model reference · open weights

AstaBrief_8B_SFT

LLMs allenai Text gen 1 build Open weights 0 dl/mo

AstaBrief_8B_SFT is an open-weight language model from allenai. AstaBrief_8B_SFT (BF16) weighs 16.4 GB; the smallest configuration that runs it is 2× RTX 3060 12 GB.

  • AstaBrief_8B_SFT is an 8-billion parameter text-generation model developed by Allen AI, designed to convert research questions and retrieved scientific literature into cited reports.
  • The model is initialized from Qwen3-8B and supports a context length of 40,960 tokens.
  • It operates in English and is released under the Apache 2.0 license.

Summary of the allenai/AstaBrief_8B_SFT model card, 2026-10-03

What it is

Released byallenai
Released2026-09-10
VRAM16.4 GB for the weights

What it runs on

Memory and cards for AstaBrief_8B_SFT (BF16)

16.4 GBweights, file size
147 MBcache per 1K tokens
753 MBruntime overhead, at least
40,960 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB … RTX 4060 Ti 16 GB
2 smaller cards
———
RTX 3090 24 GB51all 40K23.4 GB
RTX 4090 24 GB51all 40K23.4 GB
RTX 5090 32 GB112all 40K31.0 GB
L40S 48 GB225all 40K44.0 GB
A100 80 GB5012all 40K78.2 GB
H100 80 GB4611all 40K78.1 GB
RTX PRO 6000 Blackwell 96 GB5914all 40K93.8 GB
DGX Spark (GB10) 128 GB unified7117all 40K107 GB
H200 141 GB9624all 40K138 GB
B200 180 GB12731all 40K176 GB
2× RTX 3060 12 GB
tensor parallel
4135K11.6 GB a card
2× RTX 4060 Ti 16 GB
tensor parallel
102all 40K15.4 GB a card
2× RTX 4090 24 GB
tensor parallel
235all 40K23.4 GB a card
2× RTX 3090 24 GB
tensor parallel
235all 40K23.4 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
118.3 GB22.0 GB
523.2 GB41.3 GB
826.8 GB55.8 GB
1636.5 GB94.4 GB
3255.8 GB172 GB
6494.4 GB326 GB

One card, with vLLM's small-card settings.

From the model card

What allenai says about AstaBrief_8B_SFT

Read the model card

AstaBrief-8B-SFT

AstaBrief-8B-SFT is the intermediate supervised fine-tuning (SFT) checkpoint of AstaBrief-8B, a model designed to turn a research question and retrieved scientific literature excerpts into a cited report.

The model is initialized from Qwen3-8B and fine-tuned on AstaBrief_SFT_Mix, a collection of real user queries paired with reports generated from the multi-step Asta ScholarQA report generation pipeline, using a variety of backing models: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1.

For a detailed overview of the project, see our blog.

Inference and Usage

[!NOTE] Recommended prompt: This checkpoint was fine-tuned using this SFT prompt. For best results, we recommend using the same prompt format at inference time with your input query and section references. Using a different prompt or interaction format may lead to degraded or inconsistent behavior.

Following is an example inference code snippet:

from transformers import AutoModelForCausalLM, AutoTokenizer
from vllm import LLM, SamplingParams

model_name = "allenai/AstaBrief_8B_SFT"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)

# prepare the model input by providing your query and retrieved literature excerpts in our recommended prompt format
messages = [
    {"role": "user", "content": formatted_sft_prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

# instantiate sampling parameters
llm = LLM(model=model_name)
sampling_params = SamplingParams(
    temperature=0.7,
    top_p=0.95,
    max_tokens=4096,
    stop_token_ids=[tokenizer.eos_token_id],
)

# conduct text completion
outputs = llm(text, sampling_params)
output_text = outputs[0].outputs[0].text.strip()
output_ids = len(outputs[0].outputs[0].token_ids)

print("content:", output_text)

Evaluation Results

Our SFT training procedure improves the performance of the base Qwen3-8B model on the ScholarQA-CS2 test set, a set of 100 user-written computer science research questions.

ModelAverageIngredient RecallAnswer PrecisionCitation PrecisionCitation Recall
Qwen3-8B77.377.890.676.264.6
AstaBrief-8B-SFT83.785.290.487.771.3

Intended Uses and Limitations

This model is licensed under Apache 2.0 and is based on Qwen 3-8B. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Please refer to the details in our SFT dataset for more information about the data sources used in SFT training.

Training

SFT training was conducted in open-instruct on 8xH100s. Hyperparameter settings:

HyperparameterValue
#Epochs5
LR5e-06
LR SchedulerLinear
Warmup Ratio0.03
Per-device Batch Size1
Gradient Accumulation Steps4
Max sequence length32768
PrecisionBF16

Links

📝 Project blog: https://allenai.org/blog/astabrief

🤗 SFT dataset: AstaBrief_SFT_Mix

🤖 Base model: Qwen3-8B

🤖 DPO checkpoint: AstaBrief_8B

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms