Model reference · open weights

Minerva

LLMs sapienzanlp Text gen 2 builds Open weights 3k dl/mo

Minerva is an open-weight language model from sapienzanlp. Minerva-3B-base-v1.0 (BF16) weighs 11.6 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.

  • Minerva-3B-base-v1.0 is a 2.9 billion parameter text-generation model developed by Sapienza NLP, designed as a foundation model for Italian and English.
  • It is based on the Mistral architecture with a maximum context length of 16,384 tokens and was trained on 660 billion tokens from the CulturaX dataset.
  • The model is released under the Apache 2.0 license.

Summary of the sapienzanlp/Minerva-3B-base-v1.0 model card, 2026-10-01

What it is

Released bysapienzanlp
Released2024-04-19
Parameters2.9B
VRAM11.6 GB for the weights

What it runs on

Memory and cards for Minerva-3B-base-v1.0 (BF16)

11.6 GBweights, file size
651 MBruntime overhead, at least

How much memory each request adds isn't estimated yet for this architecture. The weights need at least the cards below, plus room for the context.

CardWeights alone
RTX 3060 12 GBdoes not fit
RTX 4060 Ti 16 GBfits
RTX 3090 24 GBfits
RTX 4090 24 GBfits
RTX 5090 32 GBfits
L40S 48 GBfits
A100 80 GBfits
H100 80 GBfits
RTX PRO 6000 Blackwell 96 GBfits
DGX Spark (GB10) 128 GB unifiedfits
H200 141 GBfits
B200 180 GBfits
2× RTX 3060 12 GB
split by layers (llama.cpp)
fits

Builds

Sizes, precisions & builds

BuildParamsPrecisionWeightsSmallest setup
Minerva-3B-base-v1.0 (above) ↗ 2.9BBF16 11.6 GBRTX 4060 Ti 16 GB
Minerva-350M-base-v1.0 ↗ 352MBF16 703 MBRTX 3060 12 GB

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms