Model reference · open weights

needle3

Available as managed deployment LLMs Cactus-Compute Text gen 1 variants 4k dl/mo

needle3 is an open-weight language model from Cactus-Compute. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

Released byCactus-Compute
TypeLanguage models
TaskText gen
Context8k tokens
Runs withcactus-needle
Released2026-09-16
Popularity4k downloads / month
LicenceOpen weights

About

What needle3 is

A foundation model for mobiles, wearables, robots, smart home, automotive and microcontrollers. The whole model is a single 8-29 MB file, and we trade general chat capacity to beat models 10x its size on mobile tool calls and match 2-3x bigger models on extraction.

Needle does three jobs, all of them on the device:

  • Tool calls: given the functions your app exposes, Needle picks the right ones and fills every argument from what the user said. Ask for two things and you get two calls in order; ask for something no tool covers and you get an empty list, not a guess.
  • Structured extraction: declare a shape, hand over messy text, get typed fields back: an invoice, a booking, a notification, a form. The decode grammar guarantees the output parses, and extraction generalises to classification.
  • Text embedding: the same model returns a vector for a sentence, so an app can search, match and route locally.
Read the full model card

Model

Needle 3 is a Laddered Simple Attention Network, our small-model recipe: a Monarch Hadamard MLP in place of the FFN, GQA attention with causal conv taps, engram n-gram memory read by gather, and multi-lane hyper-connections, trained so that every depth from 2 to 20 layers is a deployable model. Most of its parameters sit in the engram, so the 121M model does the arithmetic of a 50M one. The weights are compressed to CQ2-bit with Cactus Quants; a byte-level grammar compiled from your schemas constrains every token, and every response carries a calibrated confidence score from a learned head. The architecture diagram is on the release page. The repo holds the 20-layer needle3.cact, the needle3.safetensors checkpoint to fine-tune, and an engine per platform.

Benchmarks

Tool calling is exact-match accuracy on the full test splits, extraction is field micro-F1 on the full test splits.

The interactive frontier plot, the architecture and the fine-tuning results are at cactuscompute.com/needle.

Get started

pip install cactus-needle

Try it in the browser at cactuscompute.com/needle; the Python package and the source are on GitHub.

import needle

@needle.tool
def get_weather(city: str):
    "Get the current weather for a city."
    return {"city": city, "temp_c": 27, "sky": "clear"}

agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]

Every turn returns one JSON object with function_calls, the model's reasoning and a calibrated confidence; an off-topic request returns an empty list rather than a guess. The engine and the weights are fetched from this repo once and cached.

Guides

  • How to design tools for Needle 3: one tool per action, names users would say, formats in descriptions, constraints in the grammar, triggers.
  • Leveraging Needle's confidence: what the score measures, what the engine withholds, and routing on act, confirm or refuse.
  • Structured JSON extraction with Needle: the record as the only tool, typed results, classification with enums.
  • Fine-tuning Needle: the data format, the commands, reading the loss, sizing the dataset.
  • Needle Python docs: the API, the response shape, the behaviour contract, system facts, tool retrieval, offline devices, environments, the CLI.
  • What devices are supported on Needle: every platform folder, the CLI runner, the C API, the browser, WASI, air-gapped setup.
  • The .cact format: the file the engine maps and reads in place, Cactus Quants at 2.125 bits per weight, and how to parse it yourself.
  • Porting Needle 3: notes for writing your own runtime, the oracle to test against, the tensor order the container promises, the prompt on the wire, the ladder rule, retrieval with needle_embed.

Customisation

Needle was designed to be customised. Its capacity is a ladder, and a subnetwork as small as 2 layers, fine-tuned on one product's tools, runs optimally on devices far smaller than the full model needs. Fine-tuning on DroidCall lifts every subnetwork by 18 to 36 points, and from 4 layers up the tuned subnetwork passes DeepSeek V4 Flash, starting at 29M parameters.

The Python package fine-tunes with LoRA on the frozen base at the full 20 layers, then needle build [--layers N] merges the adapter, slices any subnetwork from 2 to 20 layers and exports a 4-bit .cact that runs on the same engine. The 2-bit post-training and quantisation behind the shipped model, enriched with Cactus proprietary datasets, run on the Cactus Platform.

Deploy

Every platform folder in this repo holds an engine under 1 MB that loads needle3.cact at start. needle build --platform [--layers N] fetches the engine and header and puts the weights beside them at any depth, or download the folder here:

./needle --model needle3.cact --tools tools.json --prompt "dim the living room to 30"
./needle --model needle3.cact --tools tools.json --serve

The devices guide covers the runner flags, the C API, the WASI component and air-gapped setup.

Citation

Needle 3 is built by the Cactus Compute team. If you use it in your work, please cite:

@misc{needle3_2026,
  title        = {Needle: Foundation Tool-Calling Model for Tiny Devices},
  author       = {Ndubuaku, Henry and Mosoyan, Karen and Mroz, Ja

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys needle3 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (needle3 below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"needle3","messages":[{"role":"user","content":"Hello"}]}'

Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms