Model reference · open weights
LongCat-Image-Edit is an open-weight image model from rootlocalghost. LongCat-Image-Edit-Turbo-FP8 (FP8) weighs 14.7 GB; the smallest configuration that runs it is RTX 4060 Ti 16 GB.
What it is
| Released by | rootlocalghost |
|---|---|
| Type | Image models |
| Task | Image edit |
| Runs with | transformers |
| Released | 2026-05-04 |
| Popularity | 1k downloads / month |
| Weights | 14.7 GB (LongCat-Image-Edit-Turbo-FP8 (FP8), file size) |
| Licence | Open weights |
What it runs on
Weights 14.7 GB (file size) · its biggest part 8.3 GB · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.
| Card | One 1024×1024 image | Counted memory |
|---|---|---|
| RTX 3060 12 GB | does not fit | 11.6 GB |
| RTX 4060 Ti 16 GB | tight (encoders offloaded) | 15.4 GB |
| RTX 3090 24 GB | tight | 23.4 GB |
| RTX 4090 24 GB | tight | 23.4 GB |
| RTX 5090 32 GB | fits | 31.0 GB |
| L40S 48 GB | fits | 44.0 GB |
| A100 80 GB | fits | 78.2 GB |
| H100 80 GB | fits | 78.1 GB |
| RTX PRO 6000 Blackwell 96 GB | fits | 93.8 GB |
| DGX Spark (GB10) 128 GB unified | fits | 107 GB |
| H200 141 GB | fits | 138 GB |
| B200 180 GB | fits | 176 GB |
Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more); "encoders offloaded" means only the biggest part is on the card at once — diffusers' model offload, or ComfyUI unloading the text encoder. diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.
From the model card
We introduce LongCat-Image-Edit-Turbo, the distilled version of LongCat-Image-Edit. It achieves high-quality image editing with only 8 NFEs (Number of Function Evaluations) , offering extremely low inference latency.
pip install git+https://github.com/huggingface/diffusers
[!CAUTION] 📝 Special Handling for Text Rendering
For both Text-to-Image and Image Editing tasks involving text generation, you must enclose the target text within single or double quotation marks (both English '...' / "..." and Chinese ‘...’ / “...” styles are supported).
Reasoning: The model utilizes a specialized character-level encoding strategy specifically for quoted content. Failure to use explicit quotation marks prevents this mechanism from triggering, which will severely compromise the text rendering capability.
import torch
from PIL import Image
from diffusers import LongCatImageEditPipeline
if __name__ == '__main__':
device = torch.device('cuda')
pipe = LongCatImageEditPipeline.from_pretrained("meituan-longcat/LongCat-Image-Edit-Turbo", torch_dtype= torch.bfloat16 )
# pipe.to(device, torch.bfloat16) # Uncomment for high VRAM devices (Faster inference)
pipe.enable_model_cpu_offload() # Offload to CPU to save VRAM (Required ~18 GB); slower but prevents OOM
img = Image.open('assets/test.png').convert('RGB')
prompt = '将猫变成狗'
image = pipe(
img,
prompt,
negative_prompt='',
guidance_scale=1,
num_inference_steps=8,
num_images_per_prompt=1,
generator=torch.Generator("cpu").manual_seed(43)
).images[0]
image.save('./edit_example.png')
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
Running it yourself
Rent a machine by the hour. Runs as it is with diffusers — on the machine, in Python.
# on your rented machine: pip install diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image
pipe = DiffusionPipeline.from_pretrained("rootlocalghost/LongCat-Image-Edit-Turbo-FP8", torch_dtype=torch.bfloat16).to("cuda")
start = load_image("/workspace/in.png")
image = pipe(prompt="the same scene at golden hour", image=start).images[0]
image.save("/workspace/out.png")