Model reference · open weights

ScaleCUA

Available as managed deployment LLMs OpenGVLab Vision + text 2 variants 69 dl/mo

ScaleCUA is an open-weight language model from OpenGVLab. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

MakerOpenGVLab
TypeLanguage models
TaskVision + text
Parameters (lead)8.3B
Context125k tokens
Runs withtransformers
Based onQwen/Qwen2.5-VL-7B-Instruct
Released2025-09-16
Popularity69 downloads / month
LicenceOpen weights

About

What ScaleCUA is

[📂 GitHub] [📜 Paper] [🚀 Quick Start]

Introduction

Recent advances in Vision-Language Models have enabled the development of agents capable of automating interactions with graphical user interfaces. Some computer use agents demonstrate strong performance, while they are typically built on closed-source models or inaccessible proprietary datasets. Moreover, the existing open-source datasets still remain insufficient for developing cross-platform general-purpose computer-use agents. To bridge this gap, we scale up the computer use dataset, constructed via a novel dual-loop interactive pipeline that combines an automated agent and a human expert into data collection. It spans 6 operating systems and 3 task domains, offering a large-scale and diverse corpus for training computer use agents. Building on this corpus, we develop ScaleCUA, capable of seamless operation across heterogeneous platforms. Trained on our dataset, it delivers consistent gains on several benchmarks, improving absolute success rates by +26.6 points on WebArena-Lite-v2 and +10.7 points on ScreenSpot-Pro compared to the baseline. Moreover, our ScaleCUA family achieves state-of-the-art performance across multiple benchmarks, e.g., 94.4% on MMBench-GUI L1-Hard, 60.6% on OSWorld-G and 47.4% on WebArena-Lite-v2. These results highlight the effectiveness of our data-centric methodology in scaling both GUI understanding, grounding, and cross-platform task completion. We make our data, models, and code publicly available to facilitate future research: https://github.com/OpenGVLab/ScaleCUA.


Model Loading

We provide an example code to run ScaleCUA using transformers.

from transformers import Qwen2_5_VLForConditionalGeneration, AutoTokenizer, AutoProcessor
from qwen_vl_utils import process_vision_info

# default: Load the model on the available device(s)
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    "OpenGVLab/ScaleCUA-7B", torch_dtype="auto", device_map="auto"
)

min_pixels = 3136
max_pixels = 2109744
processor = AutoProcessor.from_pretrained("OpenGVLab/ScaleCUA-7B", min_pixels=min_pixels, max_pixels=max_pixels)

Direct Action Mode as grounder

For tasks that require direct GUI grounding (e.g., identifying and clicking a specific button from a description) or serve as grounder in agentic workflow, you can use the Direct Action Mode. This mode focuses on generating immediate, executable actions based on the visual input.

  1. To enable this mode, set the system prompt as follows:
SCALECUA_SYSTEM_PROMPT_GROUNDER = '''You are an autonomous GUI agent capable of operating on desktops, mobile devices, and web browsers. Your primary function is to analyze screen captures and perform appropriate UI actions to complete assigned tasks.

## Action Space
def click(
x: float | None = None,
y: float | None = None,
clicks: int = 1,
button: str = "left",
) -> None:
"""Clicks on the screen at the specified coordinates. The `x` and `y` parameter specify where the mouse event occurs. If not provided, the current mouse position is used. The `clicks` parameter specifies how many times to click, and the `button` parameter specifies which mouse button to use ('left', 'right', or 'middle')."""
pass

def doubleClick(
x: float | None = None,
y: float | None = None,
button: str = "left",
) -> None:
"""Performs a double click. This is a wrapper function for click(x, y, 2, 'left')."""
pass

def rightClick(x: float | None = None, y: float | None = None) -> None:
"""Performs a right mouse button click. This is a wrapper function for click(x, y, 1, 'right')."""
pass

def moveTo(x: float, y: float) -> None:
"""Move the mouse to the specified coordinates."""
pass

def dragTo(
x: float | None = None, y: float | None = None, button: str = "left"
) -> None:
"""Performs a drag-to action with optional `x` and `y` coordinates and button."""
pass

def swipe(
from_coord: tuple[float, float] | None = None,
to_coord: tuple[float, float] | None = None,
direction: str = "up",
amount: float = 0.5,
) -> None:
"""Performs a swipe action on the screen. The `from_coord` and `to_coord` specify the starting and ending coordinates of the swipe. If `to_coord` is not provided, the `direction` and `amount` parameters are used to determine the swipe direction and distance. The `direction` can be 'up', 'down', 'left', or 'right', and the `amount` specifies how far to swipe relative to the screen size (0 to 1)."""
pass

def long_press(x: float, y: float, duration: int = 1) -> None:
"""Long press on the screen at the specified coordinates. The `duration` specifies how long to hold the press in seconds."""
pass

## Input Specification
- Screenshot of the current screen + task description

## Output Format
[A set of executable action command]

## Note
- Avoid action(s) that would lead to invalid states.
- The generated action(s) must exist within the defined action space.
- The generated action(s) should be enclosed within  tags.'''
  1. Use the above system prompt to generate prediction:
low_level_instruction = "Click the 'X' button in the upper right corner of the pop-up to close it and access the car selection options."

messages = [
    {
      "role": "system",
      "content":[
        {
          "type": "text",
          "text": SCALECUA_SYSTEM_PROMPT_GROUNDER,
        }
      ]
    },
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": "/path/to/your/image",
            },
            {"type": "text", "text": low_level_instruction},
        ],
    }
]

# Preparation for inference
text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
    text=[text],
    images=image_inputs,
    vide

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys scalecua for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (scalecua below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/chat/completions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"scalecua","messages":[{"role":"user","content":"Hello"}]}'

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms