Model reference · open weights

xLAM-8x22b-r

LLMs Salesforce Text gen 1 build Non-commercial 17k dl/mo

xLAM-8x22b-r is an open-weight language model from Salesforce. xLAM-8x22b-r (BF16) weighs 281 GB; the smallest configuration that runs it is 4× H100 80 GB.

xLAM-8x22b-r is a 140.6B parameter large language model developed by Salesforce for text generation and function calling. It supports a context length of 65,536 tokens and operates in English. The model is licensed under cc-by-nc-4.0 and is released exclusively for research purposes.

Summary of the Salesforce/xLAM-8x22b-r model card, 2026-10-01

What it is

Released bySalesforce
TypeLanguage models
TaskText gen
Parameters (lead)140.6B
Context65,536 tokens
Runs withtransformers
Released2024-08-28
Popularity17k downloads / month
Weights281 GB (xLAM-8x22b-r (BF16), file size)
LicenceNon-commercial

What it runs on

Memory and cards for xLAM-8x22b-r (BF16)

Weights 281 GB (file size) · KV cache 229 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 698 MB on a small card · context up to 65,536 tokens.

CardRequests at once
8K tokens each
Requests at once
32K tokens each
Longest single
request
Counted
memory
RTX 3060 12 GB … B200 180 GB
12 smaller cards
———
4× H100 80 GB
tensor parallel
82all 64K78.1 GB a card
4× A100 80 GB
tensor parallel
153all 64K78.2 GB a card
2× B200 180 GB
tensor parallel
328all 64K176 GB a card
4× RTX PRO 6000 Blackwell 96 GB
tensor parallel
4110all 64K93.8 GB a card
4× H200 141 GB
tensor parallel
13533all 64K138 GB a card
Memory needed at each load
Requests at once8K tokens each32K tokens each
1284 GB289 GB
5291 GB320 GB
8297 GB342 GB
16312 GB402 GB
32342 GB522 GB
64402 GB763 GB

On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.

Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.

From the model card

What Salesforce says about xLAM-8x22b-r

Read the model card

Welcome to the xLAM model family! Large Action Models (LAMs) are advanced large language models designed to enhance decision-making and translate user intentions into executable actions that interact with the world. LAMs autonomously plan and execute tasks to achieve specific goals, serving as the brains of AI agents. They have the potential to automate workflow processes across various domains, making them invaluable for a wide range of applications. The model release is exclusively for research purposes. A new and enhanced version of xLAM will soon be available exclusively to customers on our Platform.

Trained with ActionStudio: A Lightweight Framework for Data and Training of Action Models.

Table of Contents

Model Series

We provide a series of xLAMs in different sizes to cater to various applications, including those optimized for function-calling and general agent applications:

Model# Total ParamsContext LengthRelease DateCategoryDownload ModelDownload GGUF files
xLAM-7b-r7.24B32kSep. 5, 2024General, Function-calling🤗 Link--
xLAM-8x7b-r46.7B32kSep. 5, 2024General, Function-calling🤗 Link--
xLAM-8x22b-r141B64kSep. 5, 2024General, Function-calling🤗 Link--
xLAM-1b-fc-r1.35B16kJuly 17, 2024Function-calling🤗 Link🤗 Link
xLAM-7b-fc-r6.91B4kJuly 17, 2024Function-calling🤗 Link🤗 Link
xLAM-v0.1-r46.7B32kMar. 18, 2024General, Function-calling🤗 Link--

For our Function-calling series (more details are included at here), we also provide their quantized GGUF files for efficient deployment and execution. GGUF is a file format designed to efficiently store and load large language models, making GGUF ideal for running AI models on local devices with limited resources, enabling offline functionality and enhanced privacy.

For more details, check our GitHub and paper.

Check Latest Examples on Interaction with xLAM

Here is the latest examples and tokenizer on interacting with xLAM models.

Repository Overview

This repository is about the general tool use series. For more specialized function calling models, please take a look into our fc series here.

The instructions will guide you through the setup, usage, and integration of our model series with HuggingFace.

Framework Versions

  • Transformers 4.41.0
  • Pytorch 2.3.0+cu121
  • Datasets 2.19.1
  • Tokenizers 0.19.1

Usage

Basic Usage with Huggingface

To use the model from Huggingface, please first install the transformers library:

pip install transformers>=4.41.0

Please note that, our model works best with our provided prompt format. It allows us to extract JSON output that is similar to the function-calling mode of ChatGPT.

We use the following example to illustrate how to use our model for 1) single-turn use case, and 2) multi-turn use case

1. Single-turn use case
import json
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

torch.random.manual_seed(0)

model_name = "Salesforce/xLAM-7b-r"
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto", torch_dtype="auto", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(model_name)

# Please use our provided instruction prompt for best performance
task_instruction = """
Based on the previous context and API request history, generate an API request or a response as an AI assistant.""".strip()

format_instruction = """
The output should be of the JSON format, which specifies a list of generated function calls. The example format is as follows, please make sure the parameter type is correct. If no function call is needed, please make
tool_calls an empty list "[]".
```
{"thought": "the thought process, or an empty string", "tool_calls": [{"name": "api_name1", "arguments": {"argument1": "value1", "argument2": "value2"}}]}
```
""".strip()

# Define the input query and available tools
query = "What's the weather like in New York in fahrenheit?"

get_weather_api = {
    "name": "get_weather",
    "description": "Get the current weather for a location",
    "parameters": {
        "type": "object",
        "properties": {
            "location": {
                "type": "string",
                "description": "The city and state, e.g. San Francisco, New York"
            },
            "unit": {
                "type": "string",
                "enum": ["celsius", "fahrenheit"],
                "description": "The unit of temperature to return"

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms