Model reference · open weights

XCurOS1.2-VLBF16

LLMs XCurOS Vision + text 1 build Its own licence terms 92k dl/mo

XCurOS1.2-VLBF16 is an open-weight language model from XCurOS. XCurOS1.2-8B-VLBF16-Instruct (BF16) weighs 17.5 GB; the smallest configuration that runs it is 2× RTX 3060 12 GB.

XCurOS1.2-VLBF16 is a proprietary multimodal model developed by XCurOS for image-text-to-text tasks. It features 8.8B parameters and a context length of 262144 tokens, supporting English and Arabic. The model is designed for secure, on-premise enterprise deployments and operates under a proprietary license.

Summary of the XCurOS/XCurOS1.2-8B-VLBF16-Instruct model card, 2026-10-01

What it is

Released byXCurOS
TypeLanguage models
TaskVision + text
Parameters (lead)8.8B
Context262,144 tokens
Runs withtransformers
Released2026-02-25
Popularity92k downloads / month
Weights17.5 GB (XCurOS1.2-8B-VLBF16-Instruct (BF16), file size)
LicenceIts own licence terms

What it runs on

Memory and cards for XCurOS1.2-8B-VLBF16-Instruct (BF16)

Weights 17.5 GB (file size) · KV cache 147 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 2.0 GB on a small card · context up to 262,144 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … RTX 4060 Ti 16 GB
2 smaller cards
———
RTX 3090 24 GB3—25K23.4 GB
RTX 4090 24 GB3—25K23.4 GB
RTX 5090 32 GB9275K31.0 GB
L40S 48 GB205161K44.0 GB
A100 80 GB4812all 256K78.2 GB
H100 80 GB4411all 256K78.1 GB
RTX PRO 6000 Blackwell 96 GB5714all 256K93.8 GB
DGX Spark (GB10) 128 GB unified6917all 256K107 GB
H200 141 GB9423all 256K138 GB
B200 180 GB12531all 256K176 GB
2× RTX 3060 12 GB
tensor parallel
1—10K11.6 GB a card
2× RTX 4060 Ti 16 GB
tensor parallel
7160K15.4 GB a card
2× RTX 4090 24 GB
tensor parallel
205166K23.4 GB a card
2× RTX 3090 24 GB
tensor parallel
205166K23.4 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
120.8 GB24.4 GB
525.6 GB43.7 GB
829.2 GB58.2 GB
1638.9 GB96.9 GB
3258.2 GB174 GB
6496.9 GB329 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What XCurOS says about XCurOS1.2-VLBF16

Read the model card

XCurOS Secure Vision-Language Model

XCurOS-1.2-8B-VLBF16-Instruct is a state-of-the-art proprietary multimodal model developed exclusively by XCurOS. It provides advanced understanding and reasoning across text and visual inputs, designed for secure, enterprise-grade, and on-premise deployments.


Key Features

  • 🔐 Security-First Design Engineered to operate safely in isolated and sensitive environments, ensuring full data privacy and integrity.

  • 🧠 Advanced Multimodal Intelligence Combines text and visual perception for high-quality reasoning and understanding tasks.

  • 🖥 Agent & System Integration Ready Compatible with operating system interfaces, automation pipelines, and agent workflows.

  • 📄 Document & OCR Capabilities Extracts and interprets text from images, scanned documents, and complex layouts efficiently.

  • 🎯 Instruction-Tuned Performance Fine-tuned to execute instructions accurately and reliably.

  • ⚡ High Efficiency & Scalable Deployment Optimized for local machines and cloud infrastructure with efficient memory usage and inference speed.

  • 🌐 Long-Context & Large-Scale Reasoning Capable of handling large documents, books, and extended multi-modal datasets with coherent understanding.

  • 🧩 Extensible & Modular Architecture Easily integrated into custom applications, agent frameworks, and secure enterprise systems.


Architecture Overview

XCurOS-1.2-8B-VLBF16-Instruct architecture provides:

  • Text-vision fusion: Seamless integration of visual and textual information for robust reasoning.
  • Long-context processing: Maintains coherence across extended inputs and multi-modal datasets.
  • Stable instruction alignment: Ensures precise adherence to commands and tasks.
  • Flexible deployment: Adaptable from edge devices to cloud servers.

Use Cases

  • Secure AI assistants for sensitive enterprise data.
  • Enterprise system automation and orchestration.
  • Document understanding and knowledge extraction.
  • Visual analysis across diagrams, images, and scanned materials.
  • On-premise deployment where data privacy and intellectual property are critical.

Model Card & Metadata

  • Model Name: XCurOS-1.2-8B-VLBF16-Instruct
  • Organization: XCurOS
  • Version: 1.2
  • License: All rights reserved, proprietary software
  • Multimodal: True
  • Secure OS Ready: True
  • Intended Audience: Enterprises, research labs, and private organizations requiring secure AI solutions

Deployment Recommendations

  • Deploy locally within secure enterprise infrastructure for maximum privacy.
  • Integrate into automation pipelines or agent-based systems.
  • Ensure sufficient GPU/CPU resources for optimal performance, especially for long-context processing.

License & Usage

This is proprietary software. All rights are reserved by XCurOS. No part of this model may be copied, redistributed, or used without explicit authorization from XCurOS.


Contribution & Support

XCurOS maintains this model internally. For collaboration, licensing, or support, contact the XCurOS development team directly. All contributions, enhancements, or integrations require explicit permission from XCurOS.

Author: 35H - (f13b696767b224479d7a06bffae9a0b62e38e2e2)

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms