Model reference · open weights
reka-edge-2603 is an open-weight language model from RekaAI. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Released by | RekaAI |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 7.1B |
| Context | 16k tokens |
| Runs with | transformers |
| Released | 2026-03-11 |
| Popularity | 533 downloads / month |
| Licence | Commercial licence needed |
About
Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding, video analysis, object detection, and agentic tool-use.
Learn more about the Reka Edge in our announcement blog post.
| Benchmark | Reka Edge | Cosmos-Reason2 8B | Qwen 3.5 9B | Gemini 3 Pro |
|---|---|---|---|---|
| VQA-V2 Visual Question Answering | 88.40 | 79.82 | 83.22 | 89.78 |
| MLVU Video Understanding | 74.30 | 37.85 | 52.39 | 80.68 |
| MMVU Multimodal Video Understanding | 71.68 | 51.52 | 68.64 | 78.88 |
| RefCOCO-A Object Detection | 93.13 | 90.98 | 93.62 | 81.46 |
| RefCOCO-B Object Detection | 86.70 | 85.74 | 88.83 | 82.85 |
| VideoHallucer Hallucination | 59.57 | 51.65 | 56.00 | 66.78 |
| Mobile Actions Tool Use | 88.40 | 77.94 | 91.78 | 89.39 |
| Metric | Reka Edge | Cosmos-Reason2 8B | Qwen 3.5 9B | Gemini 3 Pro* |
|---|---|---|---|---|
| Input tokens For a 1024 x 1024 image | 331 | 1063 | 1041 | 1094 |
| End-to-end latency (in seconds) | 4.69 ± 2.48 | 10.56 ± 3.47 | 10.31 ± 1.81 | 16.67 ± 4.47 |
| TTFT (s) Time to first token | 0.522 ± 0.452 | 0.844 ± 0.923 | 0.60 ± 0.65 | 13.929 ± 3.872 |
*Gemini 3 Pro measured via API call; other models measured with local inference.
To get started:
cmake -B build
cmake --build build --target llama-server -j
cmake --build build --target llama-quantize -j
convert_reka_vlm_to_gguf.py) from the llama.cpp repo rootpython3 convert_reka_vlm_to_gguf.py /path/to/reka/weights \
--outfile /path/to/reka-text-f16.gguf \
--outtype f16
# Export the vision encoder
python3 convert_reka_vlm_to_gguf.py /path/to/reka/weights \
--mmproj \
--outfile /path/to/reka-mmproj-f16.gguf \
--outtype f16
quantize_reka_...) for simple quantizations of the model# Example usage for text decoder quantization
bash inference/hf_release/quantize_reka_q4_last8_q8.sh /path/to/reka-text-f16.gguf /path/to/reka-text-q4_last8_q8.gguf
./build/bin/llama-server -m /path/to/reka-text-f16.gguf \
--mmproj /path/to/reka-mmproj-f16.gguf \
-t 8 -c 2048 --host 0.0.0.0 --port 8080 --reasoning off \
One note: the model does not currently support reasoning, so we run llama-server with --reasoning off.
The easiest way to run the model is with the included example.py script. It uses PEP 723 inline metadata so uv resolves dependencies automatically — no manual install step:
uv run example.py --image media/hamburger.jpg --prompt "What is in this image?"
With quantization, Reka Edge can also be run on:
Reach out for support deploying Reka Edge to a custom edge compute platform.
If you prefer not to use the script, install dependencies manually and paste the code below:
uv pip install "transformers==4.57.3" torch torchvision pillow tiktoken imageio einops av
import torch
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor
model_id = "RekaAI/reka-edge-2603"
# Load processor and model
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.float16,
).eval()
# Move to MPS (Apple Silicon GPU)
device = torch.device("mps")
model = model.to(device)
# Prepare an image + text query
image_path = "media/hamburger.jpg" # included in the model repo
messages = [
{
"role": "user",
"content": [
{"type": "image", "image": image_path},
{"type": "text", "text": "What is in this image?"},
],
}
]
# Tokenize using the chat template
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
)
# Move tensors to device
for key, val in inputs.items():
if isinstance(val, torch.Tensor):
if val.is_floating_point():
inputs[key] = val.to(device=device, dtype=torch.float16)
else:
inputs[key] = val.to(device=device)
# Generate
with torch.inference_mode():
# Stop on token (end-of-turn) in addition to default EOS
sep_token_id = processor.tokenizer.convert_tokens_to_ids("")
output_ids = model.generate(
**inputs,
max_new_tokens=256,
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys reka-edge-2603 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (reka-edge-2603 below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"reka-edge-2603","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 3M free tokens every 30 days with every new account.