Model reference · open weights
xLAM-8x22b-r is an open-weight language model from Salesforce. xLAM-8x22b-r (BF16) weighs 281 GB; the smallest configuration that runs it is 4× H100 80 GB.
xLAM-8x22b-r is a 140.6B parameter large language model developed by Salesforce for text generation and function calling. It supports a context length of 65,536 tokens and operates in English. The model is licensed under cc-by-nc-4.0 and is released exclusively for research purposes.
Summary of the Salesforce/xLAM-8x22b-r model card, 2026-10-01
What it is
| Released by | Salesforce |
|---|---|
| Type | Language models |
| Task | Text gen |
| Parameters (lead) | 140.6B |
| Context | 65,536 tokens |
| Runs with | transformers |
| Released | 2024-08-28 |
| Popularity | 17k downloads / month |
| Weights | 281 GB (xLAM-8x22b-r (BF16), file size) |
| Licence | Non-commercial |
What it runs on
Weights 281 GB (file size) · KV cache 229 MB per 1,000 tokens of context, at 16 bits (vLLM's default for this build; an 8-bit cache halves it) · runtime overhead from 698 MB on a small card · context up to 65,536 tokens.
| Card | Requests at once 8K tokens each | Requests at once 32K tokens each | Longest single request | Counted memory |
|---|---|---|---|---|
| RTX 3060 12 GB … B200 180 GB 12 smaller cards | — | — | — | |
| 4× H100 80 GB tensor parallel | 8 | 2 | all 64K | 78.1 GB a card |
| 4× A100 80 GB tensor parallel | 15 | 3 | all 64K | 78.2 GB a card |
| 2× B200 180 GB tensor parallel | 32 | 8 | all 64K | 176 GB a card |
| 4× RTX PRO 6000 Blackwell 96 GB tensor parallel | 41 | 10 | all 64K | 93.8 GB a card |
| 4× H200 141 GB tensor parallel | 135 | 33 | all 64K | 138 GB a card |
| Requests at once | 8K tokens each | 32K tokens each |
|---|---|---|
| 1 | 284 GB | 289 GB |
| 5 | 291 GB | 320 GB |
| 8 | 297 GB | 342 GB |
| 16 | 312 GB | 402 GB |
| 32 | 342 GB | 522 GB |
| 64 | 402 GB | 763 GB |
On one card, with vLLM's small-card settings (2,048 tokens a step). Cards of 70 GB and more reserve more per request and more overhead — each row above uses its own card's settings.
Estimates, not measurements, checked against published vLLM startup logs. The weights are the build's file size; the cache is calculated from its config (grouped-query attention); the overhead is an estimate of vLLM's own memory with that card's default settings. "Requests at once" is how many requests of that length vLLM admits — its reservation at full length, with --max-model-len set to that length; requests that stay shorter fit more. "Longest single request" is the most one request can hold there: below the model's maximum, vLLM starts only with --max-model-len set at or under it. "Counted memory" is vLLM's default 92 % of what CUDA reports for the card (the DGX Spark: about 100 GiB of its shared 128 GB). A tensor-parallel split pools the cards' memory and speeds each token up, at the cost of the link between them; a layer split (llama.cpp) holds more but does not make one request faster. Assumes vLLM 0.10 or later.
From the model card
Welcome to the xLAM model family! Large Action Models (LAMs) are advanced large language models designed to enhance decision-making and translate user intentions into executable actions that interact with the world. LAMs autonomously plan and execute tasks to achieve specific goals, serving as the brains of AI agents. They have the potential to automate workflow processes across various domains, making them invaluable for a wide range of applications. The model release is exclusively for research purposes. A new and enhanced version of xLAM will soon be available exclusively to customers on our Platform.
Trained with ActionStudio: A Lightweight Framework for Data and Training of Action Models.
We provide a series of xLAMs in different sizes to cater to various applications, including those optimized for function-calling and general agent applications:
| Model | # Total Params | Context Length | Release Date | Category | Download Model | Download GGUF files |
|---|---|---|---|---|---|---|
| xLAM-7b-r | 7.24B | 32k | Sep. 5, 2024 | General, Function-calling | 🤗 Link | -- |
| xLAM-8x7b-r | 46.7B | 32k | Sep. 5, 2024 | General, Function-calling | 🤗 Link | -- |
| xLAM-8x22b-r | 141B | 64k | Sep. 5, 2024 | General, Function-calling | 🤗 Link | -- |
| xLAM-1b-fc-r | 1.35B | 16k | July 17, 2024 | Function-calling | 🤗 Link | 🤗 Link |
| xLAM-7b-fc-r | 6.91B | 4k | July 17, 2024 | Function-calling | 🤗 Link | 🤗 Link |
| xLAM-v0.1-r | 46.7B | 32k | Mar. 18, 2024 | General, Function-calling | 🤗 Link | -- |
For our Function-calling series (more details are included at here), we also provide their quantized GGUF files for efficient deployment and execution. GGUF is a file format designed to efficiently store and load large language models, making GGUF ideal for running AI models on local devices with limited resources, enabling offline functionality and enhanced privacy.
For more details, check our GitHub and paper.
Here is the latest examples and tokenizer on interacting with xLAM models.
This repository is about the general tool use series. For more specialized function calling models, please take a look into our fc series here.
The instructions will guide you through the setup, usage, and integration of our model series with HuggingFace.
To use the model from Huggingface, please first install the transformers library:
pip install transformers>=4.41.0
Please note that, our model works best with our provided prompt format. It allows us to extract JSON output that is similar to the function-calling mode of ChatGPT.
We use the following example to illustrate how to use our model for 1) single-turn use case, and 2) multi-turn use case
import json
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
torch.random.manual_seed(0)
model_name = "Salesforce/xLAM-7b-r"
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto", torch_dtype="auto", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(model_name)
# Please use our provided instruction prompt for best performance
task_instruction = """
Based on the previous context and API request history, generate an API request or a response as an AI assistant.""".strip()
format_instruction = """
The output should be of the JSON format, which specifies a list of generated function calls. The example format is as follows, please make sure the parameter type is correct. If no function call is needed, please make
tool_calls an empty list "[]".
```
{"thought": "the thought process, or an empty string", "tool_calls": [{"name": "api_name1", "arguments": {"argument1": "value1", "argument2": "value2"}}]}
```
""".strip()
# Define the input query and available tools
query = "What's the weather like in New York in fahrenheit?"
get_weather_api = {
"name": "get_weather",
"description": "Get the current weather for a location",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "The city and state, e.g. San Francisco, New York"
},
"unit": {
"type": "string",
"enum": ["celsius", "fahrenheit"],
"description": "The unit of temperature to return"Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.