Model reference · open weights

OmniNeural

Available as managed deployment LLMs NexaAI · community Omni (any→any) 1 variants 594 dl/mo

OmniNeural is an open-weight language model from NexaAI. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

MakerNexaAI
TypeLanguage models
TaskOmni (any→any)
Released2025-08-15
Popularity594 downloads / month
LicenceOpen weights

About

What OmniNeural is

Overview

OmniNeural is the first fully multimodal model designed specifically for Neural Processing Units (NPUs). It natively understands text, images, and audio, and runs across PCs, mobile devices, automobile, IoT, and robotics.

Demos

📱 Mobile Phone NPU - Demo on Samsung S25 Ultra

The first-ever fully local, multimodal, and conversational AI assistant that hears you and sees what you see, running natively on Snapdragon NPU for long battery life and low latency.

src="https://huggingface.co/NexaAI/OmniNeural-4B/resolve/main/assets/MOBILE_50MB.mp4" type="video/mp4">


✨ PC NPU - Capabilities Highlights

src="https://huggingface.co/NexaAI/OmniNeural-4B/resolve/main/assets/PC_demo_2_image.mov">

src="https://huggingface.co/NexaAI/OmniNeural-4B/resolve/main/assets/PC_Demo_Agent.mov">

src="https://huggingface.co/NexaAI/OmniNeural-4B/resolve/main/assets/PC_Demo_Audio.mov">


Key Features

  • Multimodal Intelligence – Processes text, image, and audio in a unified model for richer reasoning and perception.
  • NPU-Optimized Architecture – Uses ReLU ops, sparse tensors, convolutional layers, and static graph execution for maximum throughput — 20% faster than non-NPU-aware models .
  • Hardware-Aware Attention – Attention patterns tuned for NPU, lowering compute and memory demand .
  • Native Static Graph – Supports variable-length multimodal inputs with stable, predictable latency .
  • Performance Gains9× faster audio processing and 3.5× faster image processing on NPUs compared to baseline encoders .
  • Privacy-First Inference – All computation stays local: private, offline-capable, and cost-efficient.

Performance / Benchmarks

Human Evaluation (vs baselines)

  • Vision: Wins/ties in ~75% of prompts against Apple Foundation, Gemma-3n-E4B, Qwen2.5-Omni-3B.
  • Audio: Clear lead over baselines, much better than Gemma3n and Apple foundation model.
  • Text: Matches or outperforms leading multimodal baselines.

Nexa Attention Speedups

  • 9× faster audio encoding (vs Whisper encoder).
  • 3.5× faster image encoding (vs SigLIP encoder).

Architecture Overview

OmniNeural’s design is tightly coupled with NPU hardware:

  • NPU-friendly ops (ReLU > GELU/SILU).
  • Sparse + small tensor multiplications for efficiency.
  • Convolutional layers favored over linear for better NPU parallelization.
  • Hardware-aware attention patterns to cut compute cost.
  • Static graph execution for predictable latency.

Production Use Cases

  • PC & Mobile – On-device AI agents combine voice, vision, and text for natural, accurate responses.

    • Examples: Summarize slides into an email (PC)*, *extract action items from chat (mobile).
    • Benefits: Private, offline, battery-efficient.
  • Automotive – In-car assistants handle voice control, cabin safety, and environment awareness.

    • Examples: Detects risks (child unbuckled, pet left, loose objects) and road conditions (fog, construction).
    • Benefits: Decisions run locally in milliseconds.
  • IoT & Robotics – Multimodal sensing for factories, AR/VR, drones, and robots.

    • Examples: Defect detection, technician overlays, hazard spotting mid-flight, natural robot interaction.
    • Benefits: Works without network connectivity.

How to use

⚠️ Hardware requirement: OmniNeural-4B currently runs only on Qualcomm NPUs (e.g., Snapdragon-powered AIPC). Apple NPU support is planned next.

1) Install Nexa-SDK

  • Download and follow the steps under "Deploy Section" Nexa's model page: Download Windows arm64 SDK
  • (Other platforms coming soon)

2) Get an access token

Create a token in the Model Hub, then log in:

nexa config set license ''

3) Run the model

Running:

nexa infer NexaAI/OmniNeural-4B

/mic mode. Once the model is running, you can type below to record your voice directly in terminal

> /mic

For images and audio, simply drag your files into the command line. Remember to leave space between file paths.


Links & Community

  • Issues / Feedback: Use the HF Discussions tab or submit an issue in our discord or nexa-sdk github.

If you want to see more NPU-first, multimodal releases on HF, please give our model a like ❤️.

Limitation

The current model is mainly optimized for English. We will optimize other language as the next step.


Citation

@misc{
      title={OmniNeural: World’s First NPU-aware Multimodal Model},
      author={Nexa AI},
      year={2025},
      url={https://huggingface.co/NexaAI/OmniNeural-4B},
}

License

This model is released under the Creative Commons Attribution–NonCommercial 4.0 (CC BY-NC 4.0) license. Non-commercial use, modification, and redistribution are permitted with attribution. For commercial licensing, please contact dev@nexa.ai.

From the published model card. Full card on the HuggingFace links in the sidebar.

How it works

How language models work

Your prompttext / messagesTransformerattention over tokensNext-token loopgenerate + streamResponsetext · tool callsA language model reads your tokens and predicts the next one, again and again, streaming the reply back.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys omnineural for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (omnineural below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"omnineural","messages":[{"role":"user","content":"Hello"}]}'

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms