Model reference · open weights

SenseNova-U1-MoT

Available as managed deployment LLMs sensenova Omni (any→any) 1 variants 18k dl/mo

SenseNova-U1-MoT is an open-weight language model from sensenova. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released bysensenova
TypeLanguage models
TaskOmni (any→any)
Parameters (lead)17.6B
Runs withtransformers
Released2026-04-22
Popularity18k downloads / month
LicenceOpen weights

About

What SenseNova-U1-MoT is

📣 Updated News

Read the full model card
  • [2026.05.08] Add GGUF quantized checkpoints and layer-offload VRAM modes for low-VRAM single-GPU inference. See Memory-efficient inference. GGUF weights for SenseNova-U1-8B-MoT-Merger are available at 🤗 smthem/SenseNova-U1-8B-MoT-Merger-gguf — many thanks to @smthem for contributing the quantized weights.

  • [2026.05.06] Release SenseNova-U1-8B-MoT-LoRA-8step-V1.0. Please see the example script.

  • [2026.04.30] Release the preview version of the 8-step inference model SenseNova-U1-8B-MoT-8step-preview. In most cases, the image generation quality of this model closely matches that of the base model (see comparison and existing issues). To test this model, you can use the inference scripts, but with the following parameters: --cfg_scale 1.0 --num_steps 8 .

  • [2026.04.27] Initial release of the weights for SenseNova-U1-8B-MoT-SFT and SenseNova-U1-8B-MoT.

  • [2026.04.27] Initial release of the inference code for SenseNova-U1.

  • 🌟 Overview

    🚀 SenseNova U1 is a new series of native multimodal models that unifies multimodal understanding, reasoning, and generation within a monolithic architecture. It marks a fundamental paradigm shift in multimodal AI: from modality integration to true unification. Rather than relying on adapters to translate between modalities, SenseNova U1 models think-and-act across language and vision natively.

    Unifying visual understanding and generation in an end-to-end architecture from pixel to word opens tremendous possibilities, enabling highly efficient and strong understanding, generation, and interleaved reasoning in a natively multimodal manner.

    🏗️ Key Pillars:

    At the core of SenseNova U1 is NEO-unify, a novel architecture designed from the first principles for multimodal AI: It eliminates both Visual Encoder (VE) and Variational Auto-Encoder (VAE) where pixel-word information are inherently and deeply correlated. Several important features are as follows:

    • 🔗 Model language and visual information end-to-end as a unified compound.
    • 🖼️ Preserve semantic richness while maintaining pixel-level visual fidelity.
    • 🧠 Reason across modalities with high efficiency & minimal conflict via native MoTs.
    What This Unlocks:

    Powered by this new core architecture, SenseNova U1 delivers exceptional efficiency in multimodal learning:

    Left: Generation Latency vs. Averaging Performance on OneIG (EN, ZH), LongText (EN, ZH), BizGenEval (Easy, Hard), CVTG and IGenBench.
    Right: Generation Latency vs. Averaging Performance on Infographic Benchmarks, i.e., BizGenEval (Easy, Hard), and IGenBench.
    
    • 🏆 Open-source SoTA in both understanding and generation: SenseNova U1 sets a new standard for unified multimodal understanding and generation, achieving state-of-the-art performance among open-source models across a wide range of understanding, reasoning, and generation benchmarks.

    • 📖 Native interleaved image-text generation: SenseNova U1 can generate coherent interleaved text and images in a single flow with one model, enabling use cases such as practical guides and travel diaries that combine clear communication with vivid storytelling and transform complex information into intuitive visuals.

    • 📰 High-density information rendering: SenseNova U1 demonstrates strong capabilities in dense visual communication, generating richly structured layouts for knowledge illustrations, posters, presentations, comics, resumes, and other information-rich formats.

    🌍 Beyond Multimodality:
    • 🤖 Vision–Language–Action (VLA)
    • 🌐 World Modeling (WM)

    🦁 Models

    In this release, we are open-sourcing the SenseNova U1 Lite series in two sizes:

    • SenseNova U1-8B-MoT — dense backbone
    • SenseNova U1-A3B-MoT — MoE backbone
    ModelParamsHF Weights
    SenseNova-U1-8B-MoT-Infographic8B MoT🤗 link
    SenseNova-U1-8B-MoT-SFT8B MoT🤗 link
    SenseNova-U1-8B-MoT8B MoT🤗 link
    SenseNova-U1-8B-MoT-LoRA-8step-V1.00.4B🤗 link
    SenseNova-U1-A3B-MoT-SFTA3B MoT🤗 link
    SenseNova-U1-A3B-MoTA3B MoT🤗 link

    Here SFT models (×32 downsampling ratio) are trained via Und

    From the published model card. Full card on the HuggingFace links in the sidebar.

    Using it via the API

    Call it like any OpenAI endpoint

    Once AxForge deploys sensenova-u1-mot for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (sensenova-u1-mot below is illustrative; you get the exact model name on deployment.)

    $ curl -sS https://api.axforge.ai/v1/chat/completions \
      -H "Authorization: Bearer $AXFORGE_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"model":"sensenova-u1-mot","messages":[{"role":"user","content":"Hello"}]}'

    Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

    © 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms