Model reference · open weights

MolmoPoint

LLMs allenai Vision + text 1 build Open weights 7k dl/mo

MolmoPoint is an open-weight language model from allenai. MolmoPoint-8B (FP32) weighs 17.4 GB; the smallest configuration that runs it is 2× RTX 3060 12 GB.

  • MolmoPoint-8B is an 8.7B parameter vision-language model developed by the Allen Institute for AI that supports image, video, and multi-image understanding and grounding.
  • It features a new pointing mechanism for image pointing, video pointing, and video tracking, and operates with a context length of 37,376 tokens.
  • The model is trained for English and released under the Apache 2.0 license.

Summary of the allenai/MolmoPoint-8B model card, 2026-10-04

What it is

Released byallenai
Released2026-03-16
Parameters8.7B
VRAM17.4 GB for the weights

What it runs on

Memory and cards for MolmoPoint-8B (FP32)

17.4 GBweights, file size
147 MBcache per 1K tokens
753 MBruntime overhead, at least
37,376 tokenscontext max
CardRequests at onceContext maxMemory
8K each32K each
RTX 3060 12 GB … RTX 4060 Ti 16 GB
2 smaller cards
———
RTX 3090 24 GB4135K23.4 GB
RTX 4090 24 GB4134K23.4 GB
RTX 5090 32 GB102all 36K31.0 GB
L40S 48 GB215all 36K44.0 GB
A100 80 GB4912all 36K78.2 GB
H100 80 GB4611all 36K78.1 GB
RTX PRO 6000 Blackwell 96 GB5914all 36K93.8 GB
DGX Spark (GB10) 128 GB unified7017all 36K107 GB
H200 141 GB9523all 36K138 GB
B200 180 GB12631all 36K176 GB
2× RTX 3060 12 GB
tensor parallel
3—28K11.6 GB a card
2× RTX 4060 Ti 16 GB
tensor parallel
92all 36K15.4 GB a card
2× RTX 4090 24 GB
tensor parallel
235all 36K23.4 GB a card
2× RTX 3090 24 GB
tensor parallel
235all 36K23.4 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
119.3 GB22.9 GB
524.1 GB42.3 GB
827.8 GB56.8 GB
1637.4 GB95.4 GB
3256.8 GB173 GB
6495.4 GB327 GB

One card, with vLLM's small-card settings.

From the model card

What allenai says about MolmoPoint

Read the model card

MolmoPoint-8B is a fully-open VLM developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding. It has new pointing mechansim that improves image pointing, video pointing, and video tracking, see our technical report for details.

Note the huggingface MolmoPoint model does not support training, see our github repo for the training code.

Quick links:

Quick Start

Setup Conda Environment

conda create --name transformers4571 python=3.11
conda activate transformers4571
pip install transformers==4.57.1
pip install torch pillow einops torchvision accelerate decord2

Inference

We recommend running MolmoPoint with logits_processor=model.build_logit_processor_from_inputs(model_inputs) to enforce points tokens are generated in a valid way.

In MolmoPoint, instead of coordinates points will be generated as a series of special tokens, decoding the tokens back into points requires some additional metadata from the preprocessor. The metadata is returned by the preprocessor using the return_pointing_metadata flag. Then model.extract_image_points and model.extract_video_points do the decoding, they return a list of ({image_id|timestamps}, object_id, pixel_x, pixel_y) output points.

Image Pointing Example:

from transformers import AutoProcessor, AutoModelForImageTextToText
import torch
import numpy as np

checkpoint_dir = "allenai/MolmoPoint-8B"  # or path to a converted HF checkpoint

model = AutoModelForImageTextToText.from_pretrained(
    checkpoint_dir,
    trust_remote_code=True,
    dtype="auto",
    device_map="auto",
)

processor = AutoProcessor.from_pretrained(
    checkpoint_dir,
    trust_remote_code=True,
    padding_side="left",
)

image_messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Point to the boats"},
            {"type": "image", "image": "https://assets.thesparksite.com/uploads/sites/5550/2025/01/aerial-view-of-boats-yachts-water-bike-and-woode-2023-11-27-04-51-17-utc.jpg"},
            {"type": "image", "image": "https://storage.googleapis.com/ai2-playground-molmo/promptTemplates/Stock_278013497.jpeg"},
        ]
    }
]

inputs = processor.apply_chat_template(
    image_messages,
    tokenize=True,
    add_generation_prompt=True,
    return_tensors="pt",
    return_dict=True,
    padding=True,
    return_pointing_metadata=True
)
metadata = inputs.pop("metadata")
inputs = {k: v.to("cuda") for k, v in inputs.items()}

with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
    output = model.generate(
        **inputs,
        logits_processor=model.build_logit_processor_from_inputs(inputs),
        max_new_tokens=200
    )

generated_tokens = output[:, inputs["input_ids"].size(1):]
generated_text = processor.post_process_image_text_to_text(generated_tokens, skip_special_tokens=False, clean_up_tokenization_spaces=False)[0]
points = model.extract_image_points(
    generated_text,
    metadata["token_pooling"],
    metadata["subpatch_mapping"],
    metadata["image_sizes"]
)

# points as a list of [object_id, image_num, x, y]
# For multiple images, `image_num` is the index of the image the point is in
print(np.array(points))

Video Pointing Example:

video_path = "https://storage.googleapis.com/oe-training-public/demo_videos/many_penguins.mp4"
video_messages = [
    {
        "role": "user",
        "content": [
            dict(type="text", text="Point to the penguins"),
            dict(type="video", video=video_path),
        ]
    }
]

inputs = processor.apply_chat_template(
    video_messages,
    tokenize=True,
    add_generation_prompt=True,
    return_tensors="pt",
    return_dict=True,
    padding=True,
    return_pointing_metadata=True
)
metadata = inputs.pop("metadata")
inputs = {k: v.to("cuda") for k, v in inputs.items()}

with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
    output = model.generate(
        **inputs,
        logits_processor=model.build_logit_processor_from_inputs(inputs),
        max_new_tokens=200
    )

    generated_tokens = output[:, inputs['input_ids'].size(1):]
    generated_text = processor.post_process_image_text_to_text(generated_tokens, skip_special_tokens=False, clean_up_tokenization_spaces=False)[0]
    video_points = model.extract_video_points(
        generated_text,
        metadata["token_pooling"],
        metadata["subpatch_mapping"],
        metadata["timestamps"],
        metadata["video_size"]
    )

    # points as a list of [object_id, image_num, x, y]
    # For tracking, object_id uniquely identifies objects that might appear multiple frames.
    print(np.array(video_points))

License and Use

This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with Ai2’s Responsible Use Guidelines. This model is trained on third party datasets that are subject to academic and non-commercial research use only. Please review the sources to determine if this model is appropriate for your use case.

Quoted from the model card on Hugging Face. The full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms