Model reference · open weights
Molmo-0924 is an open-weight language model from allenai. Molmo-72B-0924 (FP32) weighs 147 GB; the smallest configuration that runs it is 4× L40S 48 GB.
Summary of the allenai/Molmo-72B-0924 model card, 2026-10-05
What it is
| Released by | allenai |
|---|---|
| Released | 2024-09-25 |
| Parameters | 73.3B |
| VRAM | 147 GB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory |
|---|---|---|---|
| 4K each | |||
| RTX 3060 12 GB … H200 141 GB 11 smaller cards | — | — | |
| B200 180 GB | 15 | all 4K | 176 GB |
| 4× L40S 48 GB tensor parallel | 18 | all 4K | 44.0 GB a card |
| 2× RTX PRO 6000 Blackwell 96 GB tensor parallel | 20 | all 4K | 93.8 GB a card |
| 2× H200 141 GB tensor parallel | 86 | all 4K | 138 GB a card |
| 4× H100 80 GB tensor parallel | 103 | all 4K | 78.1 GB a card |
| 4× A100 80 GB tensor parallel | 120 | all 4K | 78.2 GB a card |
| 2× B200 180 GB tensor parallel | 139 | all 4K | 176 GB a card |
| Requests at once | 4K tokens each |
|---|---|
| 1 | 149 GB |
| 5 | 155 GB |
| 8 | 159 GB |
| 16 | 169 GB |
| 32 | 191 GB |
| 64 | 234 GB |
One card, with vLLM's small-card settings.
From the model card
Molmo is a family of open vision-language models developed by the Allen Institute for AI. Molmo models are trained on PixMo, a dataset of 1 million, highly-curated image-text pairs. It has state-of-the-art performance among multimodal models with a similar size while being fully open-source. You can find all models in the Molmo family here. Learn more about the Molmo family in our announcement blog post or the paper.
Molmo 72B is based on Qwen2-72B and uses OpenAI CLIP as vision backbone. Molmo-72B achieves the highest academic benchmark score and ranks second on human evaluation, just slightly behind GPT-4o.
This checkpoint is a preview of the Molmo release. All artifacts used in creating Molmo (PixMo dataset, training code, evaluations, intermediate checkpoints) will be made available at a later date, furthering our commitment to open-source AI development and reproducibility.
Sign up here to be the first to know when artifacts are released.
Quick links:
To run Molmo, first install dependencies:
pip install einops torchvision
Then, follow these steps:
from transformers import AutoModelForCausalLM, AutoProcessor, GenerationConfig
from PIL import Image
import requests
import torch
# load the processor
processor = AutoProcessor.from_pretrained(
'allenai/Molmo-72B-0924',
trust_remote_code=True,
torch_dtype='auto',
device_map='auto'
)
# load the model
model = AutoModelForCausalLM.from_pretrained(
'allenai/Molmo-72B-0924',
trust_remote_code=True,
torch_dtype='auto',
device_map='auto'
)
# process the image and text
inputs = processor.process(
images=[Image.open(requests.get("https://picsum.photos/id/237/536/354", stream=True).raw)],
text="Describe this image."
)
# move inputs to the correct device and make a batch of size 1
inputs = {k: v.to(model.device).unsqueeze(0) for k, v in inputs.items()}
# generate output; maximum 200 new tokens; stop generation when is generated
output = model.generate_from_batch(
inputs,
GenerationConfig(max_new_tokens=200, stop_strings=""),
tokenizer=processor.tokenizer
)
# only get generated tokens; decode them to text
generated_tokens = output[0,inputs['input_ids'].size(1):]
generated_text = processor.tokenizer.decode(generated_tokens, skip_special_tokens=True)
# print the generated text
print(generated_text)
# >>> This image features an adorable black Labrador puppy sitting on a wooden deck.
# The puppy is positioned in the center of the frame, looking up at the camera...
To make inference more efficient, run with autocast:
with torch.autocast(device_type="cuda", enabled=True, dtype=torch.bfloat16):
output = model.generate_from_batch(
inputs,
GenerationConfig(max_new_tokens=200, stop_strings=""),
tokenizer=processor.tokenizer
)
We did most of our evaluation in this setting (autocast on, but float32 weights)
To even further reduce the memory requirements, the model can be run with bfloat16 weights:
model.to(dtype=torch.bfloat16)
inputs["images"] = inputs["images"].to(torch.bfloat16)
output = model.generate_from_batch(
inputs,
GenerationConfig(max_new_tokens=200, stop_strings=""),
tokenizer=processor.tokenizer
)
Note that we have observed that this can change the output of the model compared to running with float32 weights.
Molmo is supported in vLLM, however please use version <=0.7.2 until a prepreprocessing bug is fixed.
| Model | Average Score on 11 Academic Benchmarks | Human Preference Elo Rating |
|---|---|---|
| Molmo 72B (this model) | 81.2 | 1077 |
| Molmo 7B-D | 77.3 | 1056 |
| Molmo 7B-O | 74.6 | 1051 |
| MolmoE 1B | 68.6 | 1032 |
| GPT-4o | 78.5 | 1079 |
| GPT-4V | 71.1 | 1041 |
| Gemini 1.5 Pro | 78.3 | 1074 |
| Gemini 1.5 Flash | 75.1 | 1054 |
| Claude 3.5 Sonnet | 76.7 | 1069 |
| Claude 3 Opus | 66.4 | 971 |
| Claude 3 Haiku | 65.3 | 999 |
| Qwen VL2 72B | 79.4 | 1037 |
| Qwen VL2 7B | 73.7 | 1025 |
| Intern VL2 LLAMA 76B | 77.1 | 1018 |
| Intern VL2 8B | 69.4 | 953 |
| Pixtral 12B | 69.5 |
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.