Model reference · open weights
q-sit-mini is an open-weight language model from zhangzicheng. q-sit-mini (FP16) weighs 1.8 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | zhangzicheng |
|---|---|
| Type | Language models |
| Task | Image→text |
| Parameters (lead) | 894M |
| Runs with | transformers |
| Released | 2025-03-11 |
| Popularity | 9k downloads / month |
| Weights | 1.8 GB (q-sit-mini (FP16), file size) |
| Licence | Open weights |
What it runs on
Weights 1.8 GB (file size) · KV cache 12 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 1.9 GB on a small card.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB | 78 | 19 | 128K | 11.6 GB |
| RTX 4060 Ti 16 GB | 116 | 29 | 128K | 15.4 GB |
| RTX 3090 24 GB | 195 | 48 | 128K | 23.4 GB |
| RTX 4090 24 GB | 195 | 48 | 128K | 23.4 GB |
| RTX 5090 32 GB | 271 | 67 | 128K | 31.0 GB |
| L40S 48 GB | 399 | 99 | 128K | 44.0 GB |
| A100 80 GB | 739 | 184 | 128K | 78.2 GB |
| H100 80 GB | 699 | 174 | 128K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 855 | 213 | 128K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 989 | 247 | 128K | 107 GB |
| H200 141 GB | 1000+ | 323 | 128K | 138 GB |
| B200 180 GB | 1000+ | 417 | 128K | 176 GB |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 3.8 GB | 4.1 GB |
| 5 | 4.2 GB | 5.7 GB |
| 8 | 4.5 GB | 6.9 GB |
| 16 | 5.3 GB | 10.2 GB |
| 32 | 6.9 GB | 16.6 GB |
| 64 | 10.2 GB | 29.5 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). Assumes vLLM 0.10 or later.
From the model card
Q-SiT is a model for image quality scoring and interpretation. It uses a Large Language Model to perform both tasks simultaneously, recognizing the inherent connection between perception and decision-making in the human visual system. Unlike previous approaches which treat scoring and interpreting as separate tasks, Q-SiT provides a unified framework.
Project page: https://github.com/Q-Future/Q-SiT
No need to install this GitHub repo. Ensure that you use the Transformers package version 4.45.0 (pip install transformers==4.45.0).
import requests
from PIL import Image
import torch
from transformers import AutoProcessor, LlavaOnevisionForConditionalGeneration
model_id = "zhangzicheng/q-sit-mini"
# if you want to use primary version, switch to q-sit
# model_id = "zhangzicheng/q-sit"
model = LlavaOnevisionForConditionalGeneration.from_pretrained(
model_id,
torch_dtype=torch.float16,
low_cpu_mem_usage=True,
).to(0)
processor = AutoProcessor.from_pretrained(model_id)
conversation = [
{
"role": "user",
"content": [
{"type": "text", "text": "How is the clarity of the human in this image?"},
{"type": "image"},
],
},
]
prompt = processor.apply_chat_template(conversation, add_generation_prompt=True)
raw_image = Image.open(requests.get("https://github.com/Q-Future/Q-SiT/blob/main/44009500.jpg?raw=true",stream=True).raw)
inputs = processor(images=raw_image, text=prompt, return_tensors='pt').to(0, torch.float16)
output = model.generate(**inputs, max_new_tokens=200, do_sample=False)
print(processor.decode(output[0][2:], skip_special_tokens=True).split("assistant")[-1])
# very low
import torch
import requests
from PIL import Image
from transformers import AutoProcessor, LlavaOnevisionForConditionalGeneration, AutoTokenizer
import numpy as np
def wa5(logits):
logprobs = np.array([logits["Excellent"], logits["Good"], logits["Fair"], logits["Poor"], logits["Bad"]])
probs = np.exp(logprobs) / np.sum(np.exp(logprobs))
return np.inner(probs, np.array([1, 0.75, 0.5, 0.25, 0]))
model_id = "zhangzicheng/q-sit-mini"
model = LlavaOnevisionForConditionalGeneration.from_pretrained(
model_id,
torch_dtype=torch.float16,
low_cpu_mem_usage=True,
).to(0)
processor = AutoProcessor.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
# Define rating tokens
toks = ["Excellent", "Good", "Fair", "Poor", "Bad"]
ids_ = [id_[0] for id_ in tokenizer(toks)["input_ids"]]
print("Rating token IDs:", ids_)
conversation = [
{
"role": "user",
"content": [
{"type": "text", "text": "Assume you are an image quality evaluator.
Your rating should be chosen from the following five categories: Excellent, Good, Fair, Poor, and Bad (from high to low).
How would you rate the quality of this image?"},
{"type": "image"},
],
},
]
prompt = processor.apply_chat_template(conversation, add_generation_prompt=True)
# Load image
raw_image = Image.open(requests.get("https://github.com/Q-Future/Q-SiT/blob/main/44009500.jpg?raw=true",stream=True).raw)
inputs = processor(images=raw_image, text=prompt, return_tensors='pt').to(0, torch.float16)
# Manually append the assistant prefix "The quality of this image is "
prefix_text = "The quality of this image is "
prefix_ids = tokenizer(prefix_text, return_tensors="pt")["input_ids"].to(0)
inputs["input_ids"] = torch.cat([inputs["input_ids"], prefix_ids], dim=-1)
inputs["attention_mask"] = torch.ones_like(inputs["input_ids"]) # Update attention mask
# Generate exactly one token (the rating)
output = model.generate(
**inputs,
max_new_tokens=1, # Generate only the rating token
output_logits=True,
return_dict_in_generate=True,
)
# Extract logits for the generated rating token
last_logits = output.logits[-1][0] # Shape: [vocab_size]
logits_dict = {tok: last_logits[id_].item() for tok, id_ in zip(toks, ids_)}
weighted_score = wa5(logits_dict)
print("Weighted average score:", weighted_score)
# Weighted average score: 0.045549712192942585 range from 0-1
# if you want range from 0-5, multiply 5
For dataset evaluation scripts, please refer to this directory. For training information, see the Training Q-SiT section of the GitHub repository.
If you find our work useful, please cite our paper as:
@misc{zhang2025teachinglmmsimagequality,
title={Teaching LMMs for Image Quality Scoring and Interpreting},
author={Zicheng Zhang and Haoning Wu and Ziheng Jia and Weisi Lin and Guangtao Zhai},
year={2025},
eprint={2503.09197},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2503.09197},
}
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.