Model reference · open weights
Aurora-80K is an open-weight language model from AuroraAI-Research. Aurora-80K (BF16) weighs 0 MB; the smallest configuration that runs it is RTX 3060 12 GB.
Summary of the AuroraAI-Research/Aurora-80K model card, 2026-10-01
What it is
| Released by | AuroraAI-Research |
|---|---|
| Released | 2026-08-20 |
| VRAM | 0 MB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory | |
|---|---|---|---|---|
| 8K each | 32K each | |||
| RTX 3060 12 GB | 1000+ | 666 | 128K | 11.6 GB |
| RTX 4060 Ti 16 GB | 1000+ | 892 | 128K | 15.4 GB |
| RTX 3090 24 GB | 1000+ | 1000+ | 128K | 23.4 GB |
| RTX 4090 24 GB | 1000+ | 1000+ | 128K | 23.4 GB |
| RTX 5090 32 GB | 1000+ | 1000+ | 128K | 31.0 GB |
| L40S 48 GB | 1000+ | 1000+ | 128K | 44.0 GB |
| A100 80 GB | 1000+ | 1000+ | 128K | 78.2 GB |
| H100 80 GB | 1000+ | 1000+ | 128K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 1000+ | 1000+ | 128K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 1000+ | 1000+ | 128K | 107 GB |
| H200 141 GB | 1000+ | 1000+ | 128K | 138 GB |
| B200 180 GB | 1000+ | 1000+ | 128K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 440 MB | 453 MB |
| 5 | 457 MB | 520 MB |
| 8 | 470 MB | 570 MB |
| 16 | 503 MB | 705 MB |
| 32 | 570 MB | 973 MB |
| 64 | 705 MB | 1.5 GB |
One card, with vLLM's small-card settings.
From the model card
Important; This new Aurora lineup under AuroraAI-Research, and is the replacement to the inefficient and old ones on my profile, expect more sizes coming soon.
This tiny language model has exactly 80 thousand parameters, and a large (relative to its parameter count) 4096 vocab size using a factorized vocabulary representation.
The benchmark scores are the following: Wikitext-2 BPB: 3.2902 BLiMP: 52.31% Arc-Easy: 26.05%
The model was trained on around 40M tokens of fineweb-edu filtered to an educational score of 4 and above, for 2 epochs, 80 million effective tokens.
The training hardware used was the Xiaomi 14T Pro, pinned to its 4 cortex-X4 CPU cores, the time taken was around 6 hours including the preprocessing, tokenizer training.
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.