Model reference · open weights
llava-onevision-qwen2-ov is an open-weight language model from lmms-lab. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | lmms-lab |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 8.0B |
| Context | 32k tokens |
| Runs with | transformers |
| Released | 2024-06-29 |
| Popularity | 35k downloads / month |
| Licence | Open weights |
About
Play with the model on the LLaVA OneVision Chat.
The LLaVA-OneVision models are 0.5/7/72B parameter models trained on LLaVA-OneVision, based on Qwen2 language model with a context window of 32K tokens.
The model was trained on LLaVA-OneVision Dataset and have the ability to interact with images, multi-image and videos.
Feel free to share your generations in the Community tab!
We provide the simple generation process for using our model. For more details, you could refer to Github.
# pip install git+https://github.com/LLaVA-VL/LLaVA-NeXT.git
from llava.model.builder import load_pretrained_model
from llava.mm_utils import get_model_name_from_path, process_images, tokenizer_image_token
from llava.constants import IMAGE_TOKEN_INDEX, DEFAULT_IMAGE_TOKEN, DEFAULT_IM_START_TOKEN, DEFAULT_IM_END_TOKEN, IGNORE_INDEX
from llava.conversation import conv_templates, SeparatorStyle
from PIL import Image
import requests
import copy
import torch
import sys
import warnings
warnings.filterwarnings("ignore")
pretrained = "lmms-lab/llava-onevision-qwen2-7b-ov"
model_name = "llava_qwen"
device = "cuda"
device_map = "auto"
tokenizer, model, image_processor, max_length = load_pretrained_model(pretrained, None, model_name, device_map=device_map) # Add any other thing you want to pass in llava_model_args
model.eval()
url = "https://github.com/haotian-liu/LLaVA/blob/1a91fc274d7c35a9b50b3cb29c4247ae5837ce39/images/llava_v1_5_radar.jpg?raw=true"
image = Image.open(requests.get(url, stream=True).raw)
image_tensor = process_images([image], image_processor, model.config)
image_tensor = [_image.to(dtype=torch.float16, device=device) for _image in image_tensor]
conv_template = "qwen_1_5" # Make sure you use correct chat template for different models
question = DEFAULT_IMAGE_TOKEN + "\nWhat is shown in this image?"
conv = copy.deepcopy(conv_templates[conv_template])
conv.append_message(conv.roles[0], question)
conv.append_message(conv.roles[1], None)
prompt_question = conv.get_prompt()
input_ids = tokenizer_image_token(prompt_question, tokenizer, IMAGE_TOKEN_INDEX, return_tensors="pt").unsqueeze(0).to(device)
image_sizes = [image.size]
cont = model.generate(
input_ids,
images=image_tensor,
image_sizes=image_sizes,
do_sample=False,
temperature=0,
max_new_tokens=4096,
)
text_outputs = tokenizer.batch_decode(cont, skip_special_tokens=True)
print(text_outputs)
@article{li2024llavaonevision,
title={LLaVA-OneVision},
}
From the published model card. Full card on the HuggingFace links in the sidebar.
Benchmarks
As published on the model card — the maker's own numbers, not measured by AxForge.
| Task | Dataset | Metric | Score |
|---|---|---|---|
| multimodal | AI2D | accuracy | 81.400 |
| multimodal | ChartQA | accuracy | 80 |
| multimodal | DocVQA | accuracy | 90.200 |
| multimodal | InfoVQA | accuracy | 70.700 |
| multimodal | MathVerse | accuracy | 26.200 |
| multimodal | MathVista | accuracy | 63.200 |
| multimodal | MMBench | accuracy | 80.800 |
| multimodal | MME-Perception | score | 1580 |
| multimodal | MME-Cognition | score | 418 |
| multimodal | MMMU | accuracy | 48.800 |
| multimodal | MMVet | accuracy | 57.500 |
| multimodal | MMStar | accuracy | 61.700 |
| multimodal | Seed-Bench | accuracy | 75.400 |
| multimodal | Science-QA | accuracy | 96 |
| multimodal | ImageDC | accuracy | 88.900 |
| multimodal | MMLBench | accuracy | 77.100 |
| multimodal | RealWorldQA | accuracy | 66.300 |
| multimodal | Vibe-Eval | accuracy | 51.700 |
| multimodal | LLaVA-W | accuracy | 90.700 |
| multimodal | LLaVA-Wilder | accuracy | 67.800 |
| multimodal | ActNet-QA | accuracy | 56.600 |
| multimodal | EgoSchema | accuracy | 60.100 |
| multimodal | MLVU | accuracy | 64.700 |
| multimodal | MVBench | accuracy | 56.700 |
Using it via the API
Once AxForge deploys lmms-lab-llava-onevision-qwen2-ov for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (lmms-lab-llava-onevision-qwen2-ov below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"lmms-lab-llava-onevision-qwen2-ov","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.