Model reference · open weights

cogvlm2-llama3

Available as managed deployment Licence fee LLMs zai-org Text gen 1 variants 7k dl/mo

cogvlm2-llama3 is an open-weight language model from zai-org. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Makerzai-org
TypeLanguage models
TaskText gen
Parameters (lead)19.5B
Context8k tokens
Runs withtransformers
Released2024-05-16
Popularity7k downloads / month
LicenceCommercial licence needed

About

What cogvlm2-llama3 is

👋 Wechat · 💡Online Demo · 🎈Github Page · 📑 Paper 📍Experience the larger-scale CogVLM model on the ZhipuAI Open Platform.

Model introduction

We launch a new generation of CogVLM2 series of models and open source two models built with Meta-Llama-3-8B-Instruct. Compared with the previous generation of CogVLM open source models, the CogVLM2 series of open source models have the following improvements:

  1. Significant improvements in many benchmarks such as TextVQA, DocVQA.
  2. Support 8K content length.
  3. Support image resolution up to 1344 * 1344.
  4. Provide an open source model version that supports both Chinese and English.

You can see the details of the CogVLM2 family of open source models in the table below:

Model namecogvlm2-llama3-chat-19Bcogvlm2-llama3-chinese-chat-19B
Base ModelMeta-Llama-3-8B-InstructMeta-Llama-3-8B-Instruct
LanguageEnglishChinese, English
Model size19B19B
TaskImage understanding, dialogue modelImage understanding, dialogue model
Text length8K8K
Image resolution1344 * 13441344 * 1344

Benchmark

Our open source models have achieved good results in many lists compared to the previous generation of CogVLM open source models. Its excellent performance can compete with some non-open source models, as shown in the table below:

ModelOpen SourceLLM SizeTextVQADocVQAChartQAOCRbenchVCR_EASYVCR_HARDMMMUMMVetMMBench
CogVLM1.17B69.7-68.359073.934.637.352.065.8
LLaVA-1.513B61.3--337--37.035.467.7
Mini-Gemini34B74.1-----48.059.380.6
LLaVA-NeXT-LLaMA38B-78.269.5---41.7-72.1
LLaVA-NeXT-110B110B-85.779.7---49.1-80.5
InternVL-1.520B80.690.983.872014.72.046.855.482.3
QwenVL-Plus-78.991.478.1726--51.455.767.0
Claude3-Opus--89.380.869463.8537.859.451.763.3
Gemini Pro 1.5-73.586.581.3-62.7328.158.5--
GPT-4V-78.088.478.565652.0425.856.867.775.0
CogVLM2-LLaMA38B84.292.381.075683.338.044.360.480.5
CogVLM2-LLaMA3-Chinese8B85.088.474.778079.925.142.860.578.9

All reviews were obtained without using any external OCR tools ("pixel only").

Quick Start

here is a simple example of how to use the model to chat with the CogVLM2 model. For More use case. Find in our github

import torch
from PIL import Image
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_PATH = "THUDM/cogvlm2-llama3-chat-19B"
DEVICE = 'cuda' if torch.cuda.is_available() else 'cpu'
TORCH_TYPE = torch.bfloat16 if torch.cuda.is_available() and torch.cuda.get_device_capability()[0] >= 8 else torch.float16

tokenizer = AutoTokenizer.from_pretrained(
    MODEL_PATH,
    trust_remote_code=True
)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_PATH,
    torch_dtype=TORCH_TYPE,
    trust_remote_code=True,
).to(DEVICE).eval()

text_only_template = "A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions. USER: {} ASSISTANT:"

while True:
    image_path = input("image path >>>>> ")
    if image_path == '':
        print('You did not enter image path, the following will be a plain text conversation.')
        image = None
        text_only_first_query = True
    else:
        image = Image.open(image_path).convert('RGB')

    history = []

    while True:
        query = input("Human:")
        if query == "clear":
            break

        if image is None:
            if text_only_first_query:
                query = text_only_template.format(query)
                text_only_first_query = False
            else:
                old_prompt = ''
                for _, (old_query, response) in enumerate(history):
                    old_prompt += old_query + " " + response + "\n"
                query = old_prompt + "USER: {} ASSISTANT:".format(query)
        if image is None:
            input_by_model = model.build_c

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys cogvlm2-llama3 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (cogvlm2-llama3 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"cogvlm2-llama3","messages":[{"role":"user","content":"Hello"}]}'

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms