Model reference · open weights

LongCat-Image-Edit

Image meituan-longcat Image edit 1 build Open weights 23k dl/mo

LongCat-Image-Edit is an open-weight image model from meituan-longcat. LongCat-Image-Edit-Turbo (BF16) weighs 29.3 GB; the smallest configuration that runs it is RTX 4090 24 GB.

LongCat-Image-Edit-Turbo is an image-to-image model developed by meituan-longcat that performs high-quality image editing with low inference latency. It operates using only 8 Number of Function Evaluations and supports English and Chinese languages. The model is released under the Apache-2.0 license.

Summary of the meituan-longcat/LongCat-Image-Edit-Turbo model card, 2026-10-01

What it is

Released bymeituan-longcat
TypeImage models
TaskImage edit
Runs withtransformers
Released2026-02-03
Popularity23k downloads / month
Weights29.3 GB (LongCat-Image-Edit-Turbo (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for LongCat-Image-Edit-Turbo (BF16)

Weights 29.3 GB (file size) · its biggest part 16.6 GB · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.

CardOne 1024×1024 imageCounted
memory
RTX 3060 12 GB … RTX 4060 Ti 16 GBdoes not fit
RTX 3090 24 GBtight (encoders offloaded)23.4 GB
RTX 4090 24 GBtight (encoders offloaded)23.4 GB
RTX 5090 32 GBfits (encoders offloaded)31.0 GB
L40S 48 GBfits44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more); "encoders offloaded" means only the biggest part is on the card at once — diffusers' model offload, or ComfyUI unloading the text encoder. diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What meituan-longcat says about LongCat-Image-Edit

Read the model card

Introduction

We introduce LongCat-Image-Edit-Turbo, the distilled version of LongCat-Image-Edit. It achieves high-quality image editing with only 8 NFEs (Number of Function Evaluations) , offering extremely low inference latency.

Installation

pip install git+https://github.com/huggingface/diffusers

Run Image Editing

[!CAUTION] 📝 Special Handling for Text Rendering

For both Text-to-Image and Image Editing tasks involving text generation, you must enclose the target text within single or double quotation marks (both English '...' / "..." and Chinese ‘...’ / “...” styles are supported).

Reasoning: The model utilizes a specialized character-level encoding strategy specifically for quoted content. Failure to use explicit quotation marks prevents this mechanism from triggering, which will severely compromise the text rendering capability.

import torch
from PIL import Image
from diffusers import LongCatImageEditPipeline

if __name__ == '__main__':
    device = torch.device('cuda')
    pipe = LongCatImageEditPipeline.from_pretrained("meituan-longcat/LongCat-Image-Edit-Turbo", torch_dtype= torch.bfloat16 )
    # pipe.to(device, torch.bfloat16)  # Uncomment for high VRAM devices (Faster inference)
    pipe.enable_model_cpu_offload()  # Offload to CPU to save VRAM (Required ~18 GB); slower but prevents OOM
    img = Image.open('assets/test.png').convert('RGB')
    prompt = '将猫变成狗'
    image = pipe(
        img,
        prompt,
        negative_prompt='',
        guidance_scale=1,
        num_inference_steps=8,
        num_images_per_prompt=1,
        generator=torch.Generator("cpu").manual_seed(43)
    ).images[0]
    image.save('./edit_example.png')

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour. Runs as it is with diffusers — on the machine, in Python.

# on your rented machine: pip install diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image

pipe = DiffusionPipeline.from_pretrained("meituan-longcat/LongCat-Image-Edit-Turbo", torch_dtype=torch.bfloat16).to("cuda")
start = load_image("/workspace/in.png")
image = pipe(prompt="the same scene at golden hour", image=start).images[0]
image.save("/workspace/out.png")
Renting a GPU — connect, tunnels, Python
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms