Model reference · open weights
XCurOS1.2-VLBF16 is an open-weight language model from XCurOS. XCurOS1.2-8B-VLBF16-Instruct (BF16) weighs 17.5 GB; the smallest configuration that runs it is 2× RTX 3060 12 GB.
XCurOS1.2-VLBF16 is a proprietary multimodal model developed by XCurOS for image-text-to-text tasks. It features 8.8B parameters and a context length of 262144 tokens, supporting English and Arabic. The model is designed for secure, on-premise enterprise deployments and operates under a proprietary license.
Summary of the XCurOS/XCurOS1.2-8B-VLBF16-Instruct model card, 2026-10-01
What it is
| Released by | XCurOS |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 8.8B |
| Context | 262,144 tokens |
| Runs with | transformers |
| Released | 2026-02-25 |
| Popularity | 92k downloads / month |
| Weights | 17.5 GB (XCurOS1.2-8B-VLBF16-Instruct (BF16), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 17.5 GB (file size) · KV cache 147 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 2.0 GB on a small card · context up to 262,144 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … RTX 4060 Ti 16 GB 2 smaller cards | — | — | — | |
| RTX 3090 24 GB | 3 | — | 25K | 23.4 GB |
| RTX 4090 24 GB | 3 | — | 25K | 23.4 GB |
| RTX 5090 32 GB | 9 | 2 | 75K | 31.0 GB |
| L40S 48 GB | 20 | 5 | 161K | 44.0 GB |
| A100 80 GB | 48 | 12 | all 256K | 78.2 GB |
| H100 80 GB | 44 | 11 | all 256K | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | 57 | 14 | all 256K | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | 69 | 17 | all 256K | 107 GB |
| H200 141 GB | 94 | 23 | all 256K | 138 GB |
| B200 180 GB | 125 | 31 | all 256K | 176 GB |
| 2× RTX 3060 12 GB tensor parallel | 1 | — | 10K | 11.6 GB a card |
| 2× RTX 4060 Ti 16 GB tensor parallel | 7 | 1 | 60K | 15.4 GB a card |
| 2× RTX 4090 24 GB tensor parallel | 20 | 5 | 166K | 23.4 GB a card |
| 2× RTX 3090 24 GB tensor parallel | 20 | 5 | 166K | 23.4 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 20.8 GB | 24.4 GB |
| 5 | 25.6 GB | 43.7 GB |
| 8 | 29.2 GB | 58.2 GB |
| 16 | 38.9 GB | 96.9 GB |
| 32 | 58.2 GB | 174 GB |
| 64 | 96.9 GB | 329 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
XCurOS-1.2-8B-VLBF16-Instruct is a state-of-the-art proprietary multimodal model developed exclusively by XCurOS. It provides advanced understanding and reasoning across text and visual inputs, designed for secure, enterprise-grade, and on-premise deployments.
🔐 Security-First Design Engineered to operate safely in isolated and sensitive environments, ensuring full data privacy and integrity.
🧠 Advanced Multimodal Intelligence Combines text and visual perception for high-quality reasoning and understanding tasks.
🖥 Agent & System Integration Ready Compatible with operating system interfaces, automation pipelines, and agent workflows.
📄 Document & OCR Capabilities Extracts and interprets text from images, scanned documents, and complex layouts efficiently.
🎯 Instruction-Tuned Performance Fine-tuned to execute instructions accurately and reliably.
⚡ High Efficiency & Scalable Deployment Optimized for local machines and cloud infrastructure with efficient memory usage and inference speed.
🌐 Long-Context & Large-Scale Reasoning Capable of handling large documents, books, and extended multi-modal datasets with coherent understanding.
🧩 Extensible & Modular Architecture Easily integrated into custom applications, agent frameworks, and secure enterprise systems.
XCurOS-1.2-8B-VLBF16-Instruct architecture provides:
This is proprietary software. All rights are reserved by XCurOS. No part of this model may be copied, redistributed, or used without explicit authorization from XCurOS.
XCurOS maintains this model internally. For collaboration, licensing, or support, contact the XCurOS development team directly. All contributions, enhancements, or integrations require explicit permission from XCurOS.
Author: 35H - (f13b696767b224479d7a06bffae9a0b62e38e2e2)
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.