Model reference · open weights

Molmo-0924

LLMs allenai Vision + text 1 build Open weights 10k dl/mo

Molmo-0924 is an open-weight language model from allenai. Molmo-72B-0924 (FP32) weighs 147 GB; the smallest configuration that runs it is 4× L40S 48 GB.

  • Molmo-0924 is an open vision-language model developed by the Allen Institute for AI that processes image and text inputs to generate text.
  • The model contains 73.3 billion parameters and supports a context length of 4096 tokens.
  • It is trained on the PixMo dataset and is licensed under Apache 2.0 for research and educational use.

Summary of the allenai/Molmo-72B-0924 model card, 2026-10-05

What it is

Released byallenai
Released2024-09-25
Parameters73.3B
VRAM147 GB for the weights

What it runs on

Memory and cards for Molmo-72B-0924 (FP32)

147 GBweights, file size
328 MBcache per 1K tokens
1.2 GBruntime overhead, at least
4,096 tokenscontext max
CardRequests at onceContext maxMemory
4K each
RTX 3060 12 GB … H200 141 GB
11 smaller cards
——
B200 180 GB15all 4K176 GB
4× L40S 48 GB
tensor parallel
18all 4K44.0 GB a card
2× RTX PRO 6000 Blackwell 96 GB
tensor parallel
20all 4K93.8 GB a card
2× H200 141 GB
tensor parallel
86all 4K138 GB a card
4× H100 80 GB
tensor parallel
103all 4K78.1 GB a card
4× A100 80 GB
tensor parallel
120all 4K78.2 GB a card
2× B200 180 GB
tensor parallel
139all 4K176 GB a card
Memory needed at each load
Requests at once4K tokens each
1149 GB
5155 GB
8159 GB
16169 GB
32191 GB
64234 GB

One card, with vLLM's small-card settings.

From the model card

What allenai says about Molmo-0924

Read the model card

Molmo is a family of open vision-language models developed by the Allen Institute for AI. Molmo models are trained on PixMo, a dataset of 1 million, highly-curated image-text pairs. It has state-of-the-art performance among multimodal models with a similar size while being fully open-source. You can find all models in the Molmo family here. Learn more about the Molmo family in our announcement blog post or the paper.

Molmo 72B is based on Qwen2-72B and uses OpenAI CLIP as vision backbone. Molmo-72B achieves the highest academic benchmark score and ranks second on human evaluation, just slightly behind GPT-4o.

This checkpoint is a preview of the Molmo release. All artifacts used in creating Molmo (PixMo dataset, training code, evaluations, intermediate checkpoints) will be made available at a later date, furthering our commitment to open-source AI development and reproducibility.

Sign up here to be the first to know when artifacts are released.

Quick links:

Quick Start

To run Molmo, first install dependencies:

pip install einops torchvision

Then, follow these steps:

from transformers import AutoModelForCausalLM, AutoProcessor, GenerationConfig
from PIL import Image
import requests
import torch

# load the processor
processor = AutoProcessor.from_pretrained(
    'allenai/Molmo-72B-0924',
    trust_remote_code=True,
    torch_dtype='auto',
    device_map='auto'
)

# load the model
model = AutoModelForCausalLM.from_pretrained(
    'allenai/Molmo-72B-0924',
    trust_remote_code=True,
    torch_dtype='auto',
    device_map='auto'
)

# process the image and text
inputs = processor.process(
    images=[Image.open(requests.get("https://picsum.photos/id/237/536/354", stream=True).raw)],
    text="Describe this image."
)

# move inputs to the correct device and make a batch of size 1
inputs = {k: v.to(model.device).unsqueeze(0) for k, v in inputs.items()}

# generate output; maximum 200 new tokens; stop generation when  is generated
output = model.generate_from_batch(
    inputs,
    GenerationConfig(max_new_tokens=200, stop_strings=""),
    tokenizer=processor.tokenizer
)

# only get generated tokens; decode them to text
generated_tokens = output[0,inputs['input_ids'].size(1):]
generated_text = processor.tokenizer.decode(generated_tokens, skip_special_tokens=True)

# print the generated text
print(generated_text)

# >>> This image features an adorable black Labrador puppy sitting on a wooden deck.
#     The puppy is positioned in the center of the frame, looking up at the camera...

To make inference more efficient, run with autocast:

with torch.autocast(device_type="cuda", enabled=True, dtype=torch.bfloat16):
  output = model.generate_from_batch(
      inputs,
      GenerationConfig(max_new_tokens=200, stop_strings=""),
      tokenizer=processor.tokenizer
  )

We did most of our evaluation in this setting (autocast on, but float32 weights)

To even further reduce the memory requirements, the model can be run with bfloat16 weights:

model.to(dtype=torch.bfloat16)
inputs["images"] = inputs["images"].to(torch.bfloat16)
output = model.generate_from_batch(
    inputs,
    GenerationConfig(max_new_tokens=200, stop_strings=""),
    tokenizer=processor.tokenizer
)

Note that we have observed that this can change the output of the model compared to running with float32 weights.

vLLM

Molmo is supported in vLLM, however please use version <=0.7.2 until a prepreprocessing bug is fixed.

Evaluations

ModelAverage Score on 11 Academic BenchmarksHuman Preference Elo Rating
Molmo 72B (this model)81.21077
Molmo 7B-D77.31056
Molmo 7B-O74.61051
MolmoE 1B68.61032
GPT-4o78.51079
GPT-4V71.11041
Gemini 1.5 Pro78.31074
Gemini 1.5 Flash75.11054
Claude 3.5 Sonnet76.71069
Claude 3 Opus66.4971
Claude 3 Haiku65.3999
Qwen VL2 72B79.41037
Qwen VL2 7B73.71025
Intern VL2 LLAMA 76B77.11018
Intern VL2 8B69.4953
Pixtral 12B69.5

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms