Model reference · open weights
GTA1 is an open-weight language model from Salesforce. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.
Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.
What it is
| Maker | Salesforce |
|---|---|
| Type | Language models |
| Task | Vision + text |
| Parameters (lead) | 8.3B |
| Context | 125k tokens |
| Runs with | transformers |
| Released | 2025-10-01 |
| Popularity | 220 downloads / month |
| Licence | Open weights |
About
Reinforcement learning (RL) (e.g., GRPO) helps with grounding because of its inherent objective alignment—rewarding successful clicks—rather than encouraging long textual Chain-of-Thought (CoT) reasoning. Unlike approaches that rely heavily on verbose CoT reasoning, GRPO directly incentivizes actionable and grounded responses. Based on findings from our blog, we share state-of-the-art GUI grounding models trained using GRPO.
We follow the standard evaluation protocol and benchmark our model on three challenging datasets. Our method consistently achieves the best results among all open-source model families. Below are the comparative results:
| Model | Size | Open Source | ScreenSpot-V2 | ScreenSpotPro | OSWORLD-G | OSWORLD-G-Refined |
|---|---|---|---|---|---|---|
| OpenAI CUA | — | ❌ | 87.9 | 23.4 | — | — |
| Claude 3.7 | — | ❌ | 87.6 | 27.7 | — | — |
| JEDI-7B | 7B | ✅ | 91.7 | 39.5 | 54.1 | — |
| SE-GUI | 7B | ✅ | 90.3 | 47.0 | — | — |
| UI-TARS | 7B | ✅ | 91.6 | 35.7 | 47.5 | — |
| UI-TARS-1.5* | 7B | ✅ | 89.7* | 42.0* | 52.8* | 64.2* |
| UGround-v1-7B | 7B | ✅ | — | 31.1 | — | 36.4 |
| Qwen2.5-VL-32B-Instruct | 32B | ✅ | 91.9* | 48.0 | 46.5 | 59.6* |
| UGround-v1-72B | 72B | ✅ | — | 34.5 | — | — |
| Qwen2.5-VL-72B-Instruct | 72B | ✅ | 94.00* | 53.3 | — | 62.2* |
| UI-TARS | 72B | ✅ | 90.3 | 38.1 | — | — |
| OpenCUA | 7B | ✅ | 92.3 | 50.0 | 55.3 | 68.3* |
| OpenCUA | 32B | ✅ | 93.4 | 55.3 | 59.6 | 70.2* |
| GTA1-2507 (Ours) | 7B | ✅ | 92.4 (∆ +2.7) | 50.1*(∆ +8.1)* | 55.1 (∆ +2.3) | 67.7 (∆ +3.5) |
| GTA1 (Ours) | 7B | ✅ | 93.4 (∆ +0.1) | 55.5*(∆ +5.5)* | 60.1*(∆ +4.8)* | 68.8*(∆ +0.5)* |
| GTA1 (Ours) | 32B | ✅ | 95.2 (∆ +1.8) | 63.6*(∆ +8.3)* | 65.2 (∆ +5.6) | 72.2*(∆ +2.0)* |
Note:
- Model size is indicated in billions (B) of parameters.
- A dash (—) denotes results that are currently unavailable.
- A superscript asterisk (﹡) denotes our evaluated result.
- UI-TARS-1.5 7B, OpenCUA-7B, and OpenCUA-32B are applied as our baseline models.
- ∆ indicates the performance improvement (∆) of our model compared to its baseline.
We evaluate our models on the OSWorld and OSWorld-Verified benchmarks following the standard evaluation protocol. The results demonstrate strong performance across both datasets.
| Agent Model | Step | OSWorld | OSWorld-Verified |
|---|---|---|---|
| Proprietary Models | |||
| Claude 3.7 Sonnet | 100 | 28.0 | — |
| OpenAI CUA 4o | 200 | 38.1 | — |
| UI-TARS-1.5 | 100 | 42.5 | 41.8 |
| OpenAI CUA o3 | 200 | 42.9 | — |
| Open-Source Models | |||
| Aria-UI w/ GPT-4o | 15 | 15.2 | — |
| Aguvis-72B w/ GPT-4o | 15 | 17.0 | — |
| UI-TARS-72B-SFT | 50 | 18.8 | — |
| Agent S w/ Claude-3.5-Sonnet | 15 | 20.5 | — |
| Agent S w/ GPT-4o | 15 | 20.6 | — |
| UI-TARS-72B-DPO | 15 | 22.7 | — |
| UI-TARS-72B-DPO | 50 | 24.6 | — |
| UI-TARS-1.5-7B | 100 | 26.9 | 27.4 |
| Jedi-7B w/ o3 | 100 | — | 51.0 |
| Jedi-7B w/ GPT-4o | 100 | 27.0 | — |
| Agent S2 w/ Claude-3.7-Sonnet | 50 | 34.5 | — |
| Agent S2 w/ Gemini-2.5-Pro | 50 | 41.4 | 45.8 |
| Agent S2.5 w/ o3 | 100 | — | 56.0 |
| Agent S2.5 w/ GPT-5 | 100 | — | 58.4 |
| CoAct-1 w/o3 & o4mini & OpenAI CUA 4o | 150 | — | 60.8 |
| GTA1-7B-2507 w/ o3 | 100 | 45.2 | 53.1 |
| GTA1-7B-2507 w/ GPT-5 | 100 | — | 61.0 |
| GTA1-32B w/ o3 | 100 | — | 55.4 |
| GTA1-32B w/ GPT-5 | 100 | — | 63.4 |
Note: A dash (—) indicates unavailable results.
We also evaluate our models on the WindowsAgentArena benchmark, demonstrating strong performance in Windows-specific GUI automation tasks.
| Agent Model | Step | Success Rate |
|---|---|---|
| Kimi-VL | 15 | 10.4 |
| WAA | — | 19.5 |
| Jedi w/ GPT-4o | 100 | 33.7 |
| GTA1-7B-2507 w/ o3 | 100 | 47.9 |
| GTA1-7B-2507 w/ GPT-5 | 100 | 49.2 |
| GTA1-32B w/ o3 | 100 | 51.2 |
| GTA1-32B w/ GPT-5 | 100 | 50.6 |
Note: A dash (—) indicates unavailable results.
Below is a code snippet demonstrating how to run inference using a trained model.
from transformers import AutoTokenizer, AutoImageProcessor
from transformers.models.qwen2_vl.image_processing_qwen2_vl_fast import smart_resize
from PIL import Image
from io import BytesIO
import base64
import re
from vllm import LLM, SamplingParams
instruction="click start"
image_path="example.png"
CLICK_REGEXES = [
From the published model card. Full card on the HuggingFace links in the sidebar.
Using it via the API
Once AxForge deploys gta1 for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (gta1 below is illustrative; you get the exact model name on deployment.)
$ curl -sS https://api.axforge.ai/v1/chat/completions \
-H "Authorization: Bearer $AXFORGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"gta1","messages":[{"role":"user","content":"Hello"}]}'
Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.