Model reference · open weights
MolmoPoint is an open-weight language model from allenai. MolmoPoint-8B (FP32) weighs 17.4 GB; the smallest configuration that runs it is 2× RTX 3060 12 GB.
Summary of the allenai/MolmoPoint-8B model card, 2026-10-04
What it is
| Released by | allenai |
|---|---|
| Released | 2026-03-16 |
| Parameters | 8.7B |
| VRAM | 17.4 GB for the weights |
What it runs on
| Card | Requests at once | Context max | Memory | |
|---|---|---|---|---|
| 8K each | 32K each | |||
| RTX 3060 12 GB … RTX 4060 Ti 16 GB 2 smaller cards | — | — | — | |
| RTX 3090 24 GB | 4 | 1 | 35K | 23.4 GB |
| RTX 4090 24 GB | 4 | 1 | 34K | 23.4 GB |
| RTX 5090 32 GB | 10 | 2 | all 36K | 31.0 GB |
| L40S 48 GB | 21 | 5 | all 36K | 44.0 GB |
| A100 80 GB | 49 | 12 | all 36K | 78.2 GB |
| H100 80 GB | 46 | 11 | all 36K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 59 | 14 | all 36K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 70 | 17 | all 36K | 107 GB |
| H200 141 GB | 95 | 23 | all 36K | 138 GB |
| B200 180 GB | 126 | 31 | all 36K | 176 GB |
| 2× RTX 3060 12 GB tensor parallel | 3 | — | 28K | 11.6 GB a card |
| 2× RTX 4060 Ti 16 GB tensor parallel | 9 | 2 | all 36K | 15.4 GB a card |
| 2× RTX 4090 24 GB tensor parallel | 23 | 5 | all 36K | 23.4 GB a card |
| 2× RTX 3090 24 GB tensor parallel | 23 | 5 | all 36K | 23.4 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 19.3 GB | 22.9 GB |
| 5 | 24.1 GB | 42.3 GB |
| 8 | 27.8 GB | 56.8 GB |
| 16 | 37.4 GB | 95.4 GB |
| 32 | 56.8 GB | 173 GB |
| 64 | 95.4 GB | 327 GB |
One card, with vLLM's small-card settings.
From the model card
MolmoPoint-8B is a fully-open VLM developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding. It has new pointing mechansim that improves image pointing, video pointing, and video tracking, see our technical report for details.
Note the huggingface MolmoPoint model does not support training, see our github repo for the training code.
Quick links:
conda create --name transformers4571 python=3.11
conda activate transformers4571
pip install transformers==4.57.1
pip install torch pillow einops torchvision accelerate decord2
We recommend running MolmoPoint with logits_processor=model.build_logit_processor_from_inputs(model_inputs)
to enforce points tokens are generated in a valid way.
In MolmoPoint, instead of coordinates points will be generated as a series of special
tokens, decoding the tokens back into points requires some additional
metadata from the preprocessor.
The metadata is returned by the preprocessor using the return_pointing_metadata flag.
Then model.extract_image_points and model.extract_video_points do the decoding, they
return a list of ({image_id|timestamps}, object_id, pixel_x, pixel_y) output points.
from transformers import AutoProcessor, AutoModelForImageTextToText
import torch
import numpy as np
checkpoint_dir = "allenai/MolmoPoint-8B" # or path to a converted HF checkpoint
model = AutoModelForImageTextToText.from_pretrained(
checkpoint_dir,
trust_remote_code=True,
dtype="auto",
device_map="auto",
)
processor = AutoProcessor.from_pretrained(
checkpoint_dir,
trust_remote_code=True,
padding_side="left",
)
image_messages = [
{
"role": "user",
"content": [
{"type": "text", "text": "Point to the boats"},
{"type": "image", "image": "https://assets.thesparksite.com/uploads/sites/5550/2025/01/aerial-view-of-boats-yachts-water-bike-and-woode-2023-11-27-04-51-17-utc.jpg"},
{"type": "image", "image": "https://storage.googleapis.com/ai2-playground-molmo/promptTemplates/Stock_278013497.jpeg"},
]
}
]
inputs = processor.apply_chat_template(
image_messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
padding=True,
return_pointing_metadata=True
)
metadata = inputs.pop("metadata")
inputs = {k: v.to("cuda") for k, v in inputs.items()}
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
output = model.generate(
**inputs,
logits_processor=model.build_logit_processor_from_inputs(inputs),
max_new_tokens=200
)
generated_tokens = output[:, inputs["input_ids"].size(1):]
generated_text = processor.post_process_image_text_to_text(generated_tokens, skip_special_tokens=False, clean_up_tokenization_spaces=False)[0]
points = model.extract_image_points(
generated_text,
metadata["token_pooling"],
metadata["subpatch_mapping"],
metadata["image_sizes"]
)
# points as a list of [object_id, image_num, x, y]
# For multiple images, `image_num` is the index of the image the point is in
print(np.array(points))
video_path = "https://storage.googleapis.com/oe-training-public/demo_videos/many_penguins.mp4"
video_messages = [
{
"role": "user",
"content": [
dict(type="text", text="Point to the penguins"),
dict(type="video", video=video_path),
]
}
]
inputs = processor.apply_chat_template(
video_messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
padding=True,
return_pointing_metadata=True
)
metadata = inputs.pop("metadata")
inputs = {k: v.to("cuda") for k, v in inputs.items()}
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
output = model.generate(
**inputs,
logits_processor=model.build_logit_processor_from_inputs(inputs),
max_new_tokens=200
)
generated_tokens = output[:, inputs['input_ids'].size(1):]
generated_text = processor.post_process_image_text_to_text(generated_tokens, skip_special_tokens=False, clean_up_tokenization_spaces=False)[0]
video_points = model.extract_video_points(
generated_text,
metadata["token_pooling"],
metadata["subpatch_mapping"],
metadata["timestamps"],
metadata["video_size"]
)
# points as a list of [object_id, image_num, x, y]
# For tracking, object_id uniquely identifies objects that might appear multiple frames.
print(np.array(video_points))
This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with Ai2’s Responsible Use Guidelines. This model is trained on third party datasets that are subject to academic and non-commercial research use only. Please review the sources to determine if this model is appropriate for your use case.
Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.