Model reference · open weights

Aurora-80K

LLMs AuroraAI-Research Text gen 1 build Open weights 560 dl/mo

Aurora-80K is an open-weight language model from AuroraAI-Research. Aurora-80K (BF16) weighs 0 MB; the smallest configuration that runs it is RTX 3060 12 GB.

  • Aurora-80K is a text-generation model developed by AuroraAI-Research and released under the Apache 2.0 license.
  • It supports English and was trained on 80 million effective tokens of educational data.
  • The model achieved a Wikitext-2 BPB score of 3.2902 and a BLiMP score of 52.31%.

Summary of the AuroraAI-Research/Aurora-80K model card, 2026-10-01

What it is

Released byAuroraAI-Research
Released2026-08-20
VRAM0 MB for the weights

What it runs on

Memory and cards for Aurora-80K (BF16)

0 MBweights, file size
1 MBcache per 1K tokens
436 MBruntime overhead, at least
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB1000+666128K11.6 GB
RTX 4060 Ti 16 GB1000+892128K15.4 GB
RTX 3090 24 GB1000+1000+128K23.4 GB
RTX 4090 24 GB1000+1000+128K23.4 GB
RTX 5090 32 GB1000+1000+128K31.0 GB
L40S 48 GB1000+1000+128K44.0 GB
A100 80 GB1000+1000+128K78.2 GB
H100 80 GB1000+1000+128K78.1 GB
RTX PRO 6000 Blackwell 96 GB1000+1000+128K93.8 GB
DGX Spark (GB10) 128 GB unified1000+1000+128K107 GB
H200 141 GB1000+1000+128K138 GB
B200 180 GB1000+1000+128K176 GB
Memory needed at each load
Requests at once8K tokens each32K tokens each
1440 MB453 MB
5457 MB520 MB
8470 MB570 MB
16503 MB705 MB
32570 MB973 MB
64705 MB1.5 GB

One card, with vLLM's small-card settings.

From the model card

What AuroraAI-Research says about Aurora-80K

Read the model card

Important; This new Aurora lineup under AuroraAI-Research, and is the replacement to the inefficient and old ones on my profile, expect more sizes coming soon.

This tiny language model has exactly 80 thousand parameters, and a large (relative to its parameter count) 4096 vocab size using a factorized vocabulary representation.

The benchmark scores are the following: Wikitext-2 BPB: 3.2902 BLiMP: 52.31% Arc-Easy: 26.05%

The model was trained on around 40M tokens of fineweb-edu filtered to an educational score of 4 and above, for 2 epochs, 80 million effective tokens.

The training hardware used was the Xiaomi 14T Pro, pinned to its 4 cortex-X4 CPU cores, the time taken was around 6 hours including the preprocessing, tokenizer training.

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms