Model reference · open weights

JoyAI-Image-Edit

Image jdopensource Image edit 1 build Open weights 62k dl/mo

JoyAI-Image-Edit is an open-weight image model from jdopensource. JoyAI-Image-Edit-Diffusers (BF16) weighs 50.3 GB; the smallest configuration that runs it is L40S 48 GB.

JoyAI-Image-Edit is a 16.3B parameter multimodal foundation model developed by jdopensource for instruction-guided image editing. It supports English and Chinese and is designed to perform precise modifications such as object movement, rotation, and camera control based on spatial understanding. The model is released under the Apache 2.0 license.

Summary of the jdopensource/JoyAI-Image-Edit-Diffusers model card, 2026-10-01

What it is

Released byjdopensource
TypeImage models
TaskImage edit
Parameters (lead)16.3B
Runs withdiffusers
Released2026-04-10
Popularity62k downloads / month
Weights50.3 GB (JoyAI-Image-Edit-Diffusers (BF16), file size)
LicenceOpen weights

What it runs on

Memory and cards for JoyAI-Image-Edit-Diffusers (BF16)

Weights 50.3 GB (file size) · its biggest part 32.5 GB · working memory for one 1024×1024 image about 5.0 GB · overhead about 537 MB.

CardOne 1024×1024 imageCounted
memory
RTX 3060 12 GB … RTX 5090 32 GBdoes not fit
L40S 48 GBtight (encoders offloaded)44.0 GB
A100 80 GBfits78.2 GB
H100 80 GBfits78.1 GB
RTX PRO 6000 Blackwell 96 GBfits93.8 GB
DGX Spark (GB10) 128 GB unifiedfits107 GB
H200 141 GBfits138 GB
B200 180 GBfits176 GB

Estimates, not measurements: the weights are the build's file size; one 1024×1024 image needs about 5 GB of working memory (larger images more); "encoders offloaded" means only the biggest part is on the card at once — diffusers' model offload, or ComfyUI unloading the text encoder. diffusers can also place a pipeline's parts on separate cards (device_map) — not estimated here. Counted memory is 92 % of what CUDA reports for the card.

From the model card

What jdopensource says about JoyAI-Image-Edit

Read the model card

     

🐶 JoyAI-Image-Edit

JoyAI-Image-Edit is a multimodal foundation model specialized in instruction-guided image editing. It enables precise and controllable edits by leveraging strong spatial understanding, including scene parsing, relational grounding, and instruction decomposition, allowing complex modifications to be applied accurately to specified regions.

🚀 Quick Start

Requirements: Python >= 3.10, CUDA-capable GPU

Install

Note: JoyImageEditPipeline will be included in the next official diffusers release (>0.38.0). Until then, install from source as shown above.

pip install torch transformers torchvision
pip install git+https://github.com/huggingface/diffusers.git

Running with Diffusers

import torch
from PIL import Image

from diffusers import JoyImageEditPipeline

pipeline = JoyImageEditPipeline.from_pretrained("jdopensource/JoyAI-Image-Edit-Diffusers")
pipeline.to(torch.bfloat16)
pipeline.to("cuda")
pipeline.set_progress_bar_config(disable=None)
print("pipeline loaded")

img_path = "./test_images/input.png"
prompt = "Remove the construction structure from the top of the crane."
image = Image.open(img_path).convert("RGB")

inputs = {
    "image": image,
    "prompt": prompt,
    "generator": torch.manual_seed(0),
    "num_inference_steps": 40,
    "guidance_scale": 4.0,
}

print("run pipeline...")

with torch.inference_mode():
    output = pipeline(**inputs)
    image = output.images[0]
    image.save("joyai_image_edit_output.png")
    print("image saved.")

More Usages

Spatial Editing Reference

JoyAI-Image supports three spatial editing prompt patterns: Object Move, Object Rotation, and Camera Control. For the most stable behavior, we recommend following the prompt templates below as closely as possible.

1. Object Move

Use this pattern when you want to move a target object into a specified region.

Prompt template:

Move the  into the red box and finally remove the red box.

Rules:

  • Replace `` with a clear description of the target object to be moved.
  • The red box indicates the target destination in the image.
  • The phrase "finally remove the red box" means the guidance box should not appear in the final edited result.

Example:

Move the board into the red box and finally remove the red box.
2. Object Rotation

Use this pattern when you want to rotate an object to a specific canonical view.

Prompt template:

Rotate the  to show the  side view.

Supported `` values:

  • front
  • right
  • left
  • rear
  • front right
  • front left
  • rear right
  • rear left

Rules:

  • Replace `` with a clear description of the object to rotate.
  • Replace `` with one of the supported directions above.
  • This instruction is intended to change the object orientation, while keeping the object identity and surrounding scene as consistent as possible.

Examples:

Rotate the dog to show the left side view.
3. Camera Control

Use this pattern when you want to change only the camera viewpoint while keeping the 3D scene itself unchanged.

Prompt template:

Move the camera.
- Camera rotation: Yaw {y_rotation}°, Pitch {p_rotation}°.
- Camera zoom: in/out/unchanged.
- Keep the 3D scene static; only change the viewpoint.

Rules:

  • {y_rotation} specifies the yaw rotation angle in degrees.

  • {p_rotation} specifies the pitch rotation angle in degrees.

  • Camera zoom must be one of:

    • in
    • out
    • unchanged
  • The last line is important: it explicitly tells the model to preserve the 3D scene content and geometry, and only adjust the camera viewpoint.

Examples:

Move the camera.
- Camera rotation: Yaw 0.0°, Pitch -15.0°.
- Camera zoom: unchanged.
- Keep the 3D scene static; only change the viewpoint.

License Agreement

JoyAI-Image is licensed under Apache 2.0.

☎️ We're Hiring!

We are actively hiring Research Scientists, AI Infra Engineers, and Interns to join us in building next-generation generative foundation models and bringing them into real-world applications. If you’re interested, please send your resume to: huanghaoyang.ocean@jd.com

Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.

Running it yourself

Run it on a rented GPU

Rent a machine by the hour. Runs as it is with diffusers — on the machine, in Python.

# on your rented machine: pip install diffusers transformers accelerate
import torch
from diffusers import DiffusionPipeline
from diffusers.utils import load_image

pipe = DiffusionPipeline.from_pretrained("jdopensource/JoyAI-Image-Edit-Diffusers", torch_dtype=torch.bfloat16).to("cuda")
start = load_image("/workspace/in.png")
image = pipe(prompt="the same scene at golden hour", image=start).images[0]
image.save("/workspace/out.png")
Renting a GPU — connect, tunnels, Python
© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms