Model reference · open weights

Trinity-Large-TrueBase

LLMs arcee-ai Text gen 1 build Its own licence terms 234 dl/mo

Trinity-Large-TrueBase is an open-weight language model from arcee-ai. Trinity-Large-TrueBase (BF16) weighs 797 GB; the smallest configuration that runs it is 8× H200 141 GB.

What it is

Released byarcee-ai
TypeLanguage models
TaskText gen
Parameters (lead)398.6B
Context8,192 tokens
Runs withtransformers
Released2026-01-27
Popularity234 downloads / month
Weights797 GB (Trinity-Large-TrueBase (BF16), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for Trinity-Large-TrueBase (BF16)

Weights 797 GB (file size) · KV cache 61 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · plus 1.5 GB a request for its sliding-window layers · runtime overhead from 1.5 GB on a small card · context up to 8,192 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … 8× A100 80 GB
14 smaller cards
———
8× H200 141 GB
tensor parallel
117—all 8K138 GB a card
8× B200 180 GB
tensor parallel
255—all 8K176 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
1801 GB—
5809 GB—
8815 GB—
16831 GB—
32863 GB—
64928 GB—

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (attention with sliding-window layers); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What arcee-ai says about Trinity-Large-TrueBase

src="https://cdn-uploads.huggingface.co/production/uploads/6435718aaaef013d1aec3b8b/i-v1KyAMOW_mgVGeic9WJ.png" alt="Arcee Trinity Large" style="max-width: 100%; height: auto;" >

Read the full model card

Trinity-Large-TrueBase

Introduction

Trinity-Large-TrueBase is a base pretraining checkpoint from Arcee AI's Trinity Large training run. It is a 398B-parameter sparse Mixture-of-Experts (MoE) model with approximately 13B active parameters per token. The checkpoint was captured after 10 trillion tokens of pretraining, prior to learning-rate annealing and before any instruction tuning or reinforcement learning.

This checkpoint is intended for research, probing, ablation studies, and downstream fine-tuning and comes without any pre-baked alignment, instruction formatting, or preference optimization.

More details on the training of Trinity Large are available in the technical report.

Model Variants

The Trinity Large family consists of three checkpoints from the same training run:

  • Trinity-Large-TrueBase (this release): 10T-token pre-anneal checkpoint with no instruction data
  • Trinity-Large-Thinking: Reasoning-optimized, agentic post-training with extended chain-of-thought
  • Trinity-Large-Base: Full 17T-token pretrained foundation model with mid-training anneals
  • Trinity-Large-Preview: Lightly post-trained, chat-ready model undergoing active RL

Architecture

Trinity-Large-TrueBase uses a sparse MoE configuration designed to maximize efficiency while maintaining large-scale capacity.

HyperparameterValue
Total parameters~398B
Active parameters per token~13B
Experts256
Active experts4
Routing strategy4-of-256 (1.56% sparsity)
Dense layers6
Pretraining context length8,192
ArchitectureSparse MoE (AfmoeForCausalLM)

Note: Extended context support (e.g., 512k) was introduced after this checkpoint and is not available in TrueBase.

Benchmark Results

BenchmarkN-shotMetricScoreStderr
arc_challenge_0shot0acc_norm,none0.6237±0.0142
bbh_fewshot3exact_match,remove_whitespace0.5784±0.0054
gpqa_diamond_5shot5acc_norm,none0.4091±0.0350
gpqa_diamond_generative_5shot5exact_match,flexible-extract0.3788±0.0346
gsm8k_8shot8exact_match,flexible-extract0.8036±0.0109
gsm8k_cot8exact_match,flexible-extract0.8044±0.0109
hellaswag_5shot5acc_norm,none0.8813±0.0032
humaneval_plus0pass@1,create_test0.5183±0.0391
leaderboard_math_hard4exact_match,none0.2696±0.0113
mbpp_plus3pass_at_1,none0.8095±0.0202
minerva_math5004math_verify,none0.4820±0.0224
mmlu_5shot5acc,none0.7845±0.0033
mmlu_generative_5shot5exact_match,get_response0.7848±0.0033
mmlu_pro5exact_match,custom-extract0.5160±0.0044
triviaqa_5shot5exact_match,remove_whitespace0.8096±0.0029
winogrande_5shot5acc,none0.8145±0.0109

Training Configuration

Pretraining

  • Training tokens: 10 trillion
  • Checkpoint type: Pre-anneal
  • Instruction data: None
  • RLHF or post-training: None

This checkpoint branches from the main Trinity Large run at the 10T-token mark, prior to learning-rate decay or post-training phases.

Optimizers

Optimizer learning rates after WSD warm-up:

  • Adam learning rate: 2e-4
  • Muon learning rate: 8e-4

Muon was used to support larger critical batch sizes in a highly sparse MoE regime.

Infrastructure

  • Hardware: 2,048 NVIDIA B300 GPUs
  • Parallelism: HSDP + Expert Parallelism
  • Compute partner: Prime Intellect
  • Data partner: Datology

Intended Use

  • Studying emergent behavior from large-scale pretraining
  • Sparse MoE routing and load-balancing research
  • Interpretability, probing, and ablation studies
  • Domain-specific fine-tuning from a clean base
  • Academic and industrial foundation model research

Rationale for Release

Most base model releases include instruction data, annealed training dynamics, or early alignment stages. Trinity-Large-TrueBase excludes these, providing an opportunity to study what large-scale models learn from pretraining data alone. This checkpoint is intended as a foundation for research rather than as a finished conversational assistant.

Known Limitations

  • Not aligned for safety, helpfulness, or conversational tone
  • Requires substantial compute and expertise to fine-tune
  • May exhibit raw or unstable behaviors typical of unaligned models
  • No extended-context tuning beyond the 8K pretraining window

License

Trinity-Large-TrueBase is released under the OpenMDW License, version 1.1 (OpenMDW-1.1).

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms