Model reference · open weights

q-sit-mini

LLMs zhangzicheng · community Image→text 1 build Open weights 9k dl/mo

q-sit-mini is an open-weight language model from zhangzicheng. q-sit-mini (FP16) weighs 1.8 GB; the smallest configuration that runs it is RTX 3060 12 GB.

What it is

Released byzhangzicheng
TypeLanguage models
TaskImage→text
Parameters (lead)894M
Runs withtransformers
Released2025-03-11
Popularity9k downloads / month
Weights1.8 GB (q-sit-mini (FP16), file size)
LicenceOpen weights

What it runs on

Memory and cards for q-sit-mini (FP16)

Weights 1.8 GB (file size) · KV cache 12 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 1.9 GB on a small card.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB7819128K11.6 GB
RTX 4060 Ti 16 GB11629128K15.4 GB
RTX 3090 24 GB19548128K23.4 GB
RTX 4090 24 GB19548128K23.4 GB
RTX 5090 32 GB27167128K31.0 GB
L40S 48 GB39999128K44.0 GB
A100 80 GB739184128K78.2 GB
H100 80 GB699174128K78.1 GB
RTX PRO 6000 Blackwell 96 GB855213128K93.8 GB
DGX Spark (GB10) 128 GB unified989247128K107 GB
H200 141 GB1000+323128K138 GB
B200 180 GB1000+417128K176 GB
Memory needed at each load
Requests at once8K tokens each32K tokens each
13.8 GB4.1 GB
54.2 GB5.7 GB
84.5 GB6.9 GB
165.3 GB10.2 GB
326.9 GB16.6 GB
6410.2 GB29.5 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.

From the model card

What zhangzicheng says about q-sit-mini

Q-SiT is a model for image quality scoring and interpretation. It uses a Large Language Model to perform both tasks simultaneously, recognizing the inherent connection between perception and decision-making in the human visual system. Unlike previous approaches which treat scoring and interpreting as separate tasks, Q-SiT provides a unified framework.

Project page: https://github.com/Q-Future/Q-SiT

Read the full model card

Quicker Start with Hugging Face AutoModel

No need to install this GitHub repo. Ensure that you use the Transformers package version 4.45.0 (pip install transformers==4.45.0).

Image Quality Interpreting Chat

import requests
from PIL import Image
import torch
from transformers import AutoProcessor, LlavaOnevisionForConditionalGeneration

model_id = "zhangzicheng/q-sit-mini"
# if you want to use primary version, switch to q-sit
# model_id = "zhangzicheng/q-sit"

model = LlavaOnevisionForConditionalGeneration.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    low_cpu_mem_usage=True,
).to(0)

processor = AutoProcessor.from_pretrained(model_id)

conversation = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "How is the clarity of the human in this image?"},
            {"type": "image"},
        ],
    },
]
prompt = processor.apply_chat_template(conversation, add_generation_prompt=True)

raw_image = Image.open(requests.get("https://github.com/Q-Future/Q-SiT/blob/main/44009500.jpg?raw=true",stream=True).raw)

inputs = processor(images=raw_image, text=prompt, return_tensors='pt').to(0, torch.float16)

output = model.generate(**inputs, max_new_tokens=200, do_sample=False)
print(processor.decode(output[0][2:], skip_special_tokens=True).split("assistant")[-1])
# very low

Image Quality Scoring

import torch
import requests
from PIL import Image
from transformers import AutoProcessor, LlavaOnevisionForConditionalGeneration, AutoTokenizer
import numpy as np

def wa5(logits):
    logprobs = np.array([logits["Excellent"], logits["Good"], logits["Fair"], logits["Poor"], logits["Bad"]])
    probs = np.exp(logprobs) / np.sum(np.exp(logprobs))
    return np.inner(probs, np.array([1, 0.75, 0.5, 0.25, 0]))

model_id = "zhangzicheng/q-sit-mini"
model = LlavaOnevisionForConditionalGeneration.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    low_cpu_mem_usage=True,
).to(0)

processor = AutoProcessor.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)

# Define rating tokens
toks = ["Excellent", "Good", "Fair", "Poor", "Bad"]
ids_ = [id_[0] for id_ in tokenizer(toks)["input_ids"]]
print("Rating token IDs:", ids_)

conversation = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Assume you are an image quality evaluator.
Your rating should be chosen from the following five categories: Excellent, Good, Fair, Poor, and Bad (from high to low).
How would you rate the quality of this image?"},
            {"type": "image"},
        ],
    },
]
prompt = processor.apply_chat_template(conversation, add_generation_prompt=True)

# Load image
raw_image = Image.open(requests.get("https://github.com/Q-Future/Q-SiT/blob/main/44009500.jpg?raw=true",stream=True).raw)
inputs = processor(images=raw_image, text=prompt, return_tensors='pt').to(0, torch.float16)

# Manually append the assistant prefix "The quality of this image is "
prefix_text = "The quality of this image is "
prefix_ids = tokenizer(prefix_text, return_tensors="pt")["input_ids"].to(0)
inputs["input_ids"] = torch.cat([inputs["input_ids"], prefix_ids], dim=-1)
inputs["attention_mask"] = torch.ones_like(inputs["input_ids"])  # Update attention mask

# Generate exactly one token (the rating)
output = model.generate(
    **inputs,
    max_new_tokens=1,  # Generate only the rating token
    output_logits=True,
    return_dict_in_generate=True,
)

# Extract logits for the generated rating token
last_logits = output.logits[-1][0]  # Shape: [vocab_size]
logits_dict = {tok: last_logits[id_].item() for tok, id_ in zip(toks, ids_)}
weighted_score = wa5(logits_dict)
print("Weighted average score:", weighted_score)
# Weighted average score: 0.045549712192942585  range from 0-1
# if you want range from 0-5, multiply 5

For dataset evaluation scripts, please refer to this directory. For training information, see the Training Q-SiT section of the GitHub repository.

Citation

If you find our work useful, please cite our paper as:

@misc{zhang2025teachinglmmsimagequality,
      title={Teaching LMMs for Image Quality Scoring and Interpreting},
      author={Zicheng Zhang and Haoning Wu and Ziheng Jia and Weisi Lin and Guangtao Zhai},
      year={2025},
      eprint={2503.09197},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2503.09197},
}

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms