edit_image 工具使用 Qwen-Image 2.1 图像编辑工作流。
工具参数:prompt(必填)、input_image(必填)、reference_images(≤15)、output_path、resolution、steps、seed、negative_prompt、custom_size、aspect_ratio、megapixels。
-
工作流
cfg=1,负面提示词无效。 -
Qwen-Image-2.1 原生支持 RGBA 透明图
0. 什么是否放弃使用edit_image工具
-
单纯的像素级操作不应该使用模型(剪切等)。
-
生图工作流(T2I)不应该使用 edit_image 工具。
1. 提示词指导
2.1 引用规则
- 用
<image1>、<image2>… 指代图片:<image1>就是input_image(编辑目标),其余按reference_images的列表顺序依次编号。 - 上传顺序 = 编号顺序。调换
reference_images的顺序会静默改变每个编号指向的对象。
2.2 指导
- 指令要具体、可执行:说清"改什么、改成什么、改到什么程度"。
Change the background to a sunset beach远好于"让它好看一点"。 - 属性解耦:每个子指令只解耦并修改一个属性;需要改多处时,用多个短句或并列动词短语逐项写清楚。例如:
Keep the character and pose in <image1> unchanged. Use the shirt from <image2>. Preserve facial identity and hair.。 - 需要保留的东西显式写出来:
keep the face, hairstyle and lighting unchanged/preserve her identity and the product shape。模型对人物与商品身份做了优化,但明确要求"保持不变"仍然更稳。 - 画面里的文字:要把文字原样渲染出来就用引号给出确切内容(
a neon sign that reads "QWEN IMAGE 2.1")。改字时最好同时说明"原文 → 新文",并保留引号。模板:A [surface/medium] that reads "[YOUR TEXT]" in [typography style], [context], [lighting] - 多图合成:在一条指令里说清每张图的角色和它们的关系,例如"把
<image2>的衣服穿到<image1>的人身上,背景换成<image3>的街道,保持<image1>的脸不变"。
2.3 透明背景
需要显式提示透明背景,并非自动透明,推荐使用固定句式:
This is an RGBA image with transparency. <你的描述>. The image has alpha channel and the background is transparent.
例:
This is an RGBA image with transparency. A cute cartoon dragon sticker. The image has alpha channel and the background is transparent.
-
透明输出必须保存为支持 alpha 的格式
-
想"从照片里抠出主体"也可以用同一套说法让它输出透明背景。例:
This is an RGBA image with transparency. The sneaker from <image1>, isolated on a clean edge with a soft contact shadow. The image has alpha channel and the background is transparent.
2.4 在指定的区域修改
Qwen-Image 2.1 支持通过独立掩码(separate masks) 来指定局部编辑区域。
独立掩码
- 是什么:把一张掩码当空间约束传进来——掩码区内按提示词重绘,区外逐像素保留。一个掩码参数都不给时,整张画布都能改。
- 两种给法,二选一:
mask_image(你自己给一张黑白掩码图,白 = 重绘)或rects(给矩形列表,矩形内 = 重绘)。两者本质是同一件东西——一张掩码——只是一个由你提供、一个由工具按坐标现做;同时给会被直接拒绝。逐项说明见 §3。 - 掩码图必须是黑白/灰度、且不带 alpha:工具固定按红通道取值。带 alpha 的图按它自带的 RGB 判色(ComfyUI 读图是 RGB 直取、alpha 单独拆出去),所以透明区并不等于「不重绘」——Photoshop 导出的透明区常带白色 RGB,那会被当成重绘区。
- 保护强度:实测掩码区外逐像素保留(与「什么都不改」同量级),而且提示词吃不掉它——同一张掩码配「把整张图变成夜间场景」,区外照样不变。所以提示词只需写掩码内要变成什么,不必再反复强调「其余保持不变」。
- 注意:
- 不能和
custom_size同时用:走掩码时输出尺寸改由input_image决定,两者冲突会被直接拒绝;要扩展画布请先使用pad_image对input_image进行处理。 - 掩码边界可能会留一圈细锯齿,可通过轻微扩大重绘区域的方式避免。
- 掩码图最好与原图同尺寸,否则会被拉伸到原图大小。
- 掩码不占
reference_images的位置;也不要把掩码图当参考图传——参考图是「内容」,模型只当它是另一张画,给不了任何空间约束。
- 不能和
2.5 输出尺寸
默认情况下,custom_size 关闭、resolution为 0,输出画布就是 <image1> 的宽高尺寸本身**,宽高各自对齐到 32 的倍数。
实测:输入 1208×1208 → 输出 1216×1216;输入 2368×1344 → 输出 2368×1344。
-
resolution > 0会把输入图缩到约resolution²的面积(各图保持比例、对齐 32),输出画布随之缩小。实测resolution=1024+ 输入 2368×1344 → 输出 1344×768。要input_image尺寸保持resolution为0(默认)。 -
custom_size=true意味着自定义输出尺寸,一般只有高清放大 / 指定输出尺寸时才需要打开, 开启参数后将由aspect_ratio+megapixels决定画布最终形态(megapixels基准 1 ≈ 1024×1024)。 -
扩展图片需要使用
pad_image先把input_image扩展到目标尺寸,再用mask_from_transparency(使用该工具时,expand=48 是一个表现良好的数值。) 生成掩码,再使用扩展后的图片和掩码作为input_image和mask_image传给edit_image。
2.6 如何控制reference_images中的物体在结果图片上的尺寸
工具会先把参考图统一缩放到固定大小,物体在成图里的大小,不由参考图有多少像素决定,而由物体在参考图里占了多大一块地方决定。
有效做法1:调整物体在图里占的比例,需要修改参考图。
有效做法2:在提示词里把尺寸说清楚(最省事)
- 中文:
蘑菇串约占画面高度的五分之一、蘑菇串大约有小臂那么长 - 英文:
the skewer is about as long as her forearm、occupying about one fifth of the frame height
能和做法1叠加:留白定大方向,提示词做微调。
3. 参数使用场景
| 参数 | 工作流默认 | 怎么用 |
| --- | --- | --- |
| input_image | 必填 | 编辑目标 = <image1>;默认输出画布就是它的宽高尺寸(各边 32 对齐,见 §2.5) |
| reference_images | 无 | 依次为 <image2>…;与 input_image 合计最多 16 张(工作流容量),官方模型支持 ≤10 张参考图 —— 超过 10 张属于超出官方验证范围,建议优先挑最必要的几张 |
| prompt | 必填 | 见第 2 节。省略不是选项(必填) |
| mask_image | 无 | 掩码图的绝对路径,白 = 重绘、黑 = 保护。请用不带 alpha 通道的黑白/灰度图:工具固定取红通道,带透明的图按它自带的 RGB 判色(透明区 ≠ 不重绘)。这张图会一并上传给 ComfyUI,最好与原图同尺寸,否则会被拉伸到原图大小。与 rects 互斥,只能给一个 |
| rects | 无 | 矩形列表 [{x,y,width,height}, …](像素,原点左上,四个字段都必填),矩形内 = 重绘;不用先画掩码图,工具直接按坐标算。给空数组 [] 会被拒绝。与 mask_image 互斥,只能给一个 |
| mask_invert | false | 把已经算出来的掩码整体反过来:白 = 重绘 ↔ 黑 = 重绘。mask_image 与 rects 都适用。单独传它不会让掩码生效——必须同时给一个来源 |
| steps | 25 | 生成步数,保持默认就好。 |
| resolution | 0 | 所有输入图(含参考图)会被缩放到约 resolution × resolution 像素,各自保持比例、对齐到 32 的倍数。0 = 不缩放,每张按自己尺寸取整,输出画布 = <image1> 尺寸。取值 0–4096。⚠️ 它还会改掉画布尺寸:> 0 时输出画布跟着缩到约 resolution² 面积(实测 1024 + 输入 2368×1344 → 输出 1344×768)——要原尺寸就留 0。|
| custom_size | 关 | 关着(默认)= 输出画布用 <image1> 的宽高尺寸(各边 32 对齐),普通改图保持关闭即可,成图不会比原图小。要做高清放大 / 指定输出尺寸才打开它,此时画布改用 aspect_ratio + megapixels 计算 |
| aspect_ratio / megapixels | 1:1 / 1 | 仅在 custom_size 打开时才参与算画布;aspect_ratio 只给比例,真正的像素量由 megapixels 决定(基准 1 ≈ 1024×1024),宽高各自按 32 对齐,因此比例是逼近、尺寸不精确。 |
| negative_prompt | 空 | 省略 = 空,不会沿用工作流里残留的文字。工作流里 cfg=1,此时负向提示词不生效——想让它生效得先把工作流 cfg 调高 |
| seed | 随机 | 复现同一张图用同一个值 |
补充:工作流采样器 euler + simple 调度、cfg=1、denoise=1,出图节点 SaveImageAdvanced。
4. 语言偏好
- 中英文都可以:官方 README 与 Diffusers 示例用英文,Qwen 自己的中文文档与中文示例同样是一等公民;这是个中文团队的模型,中文指令不会因为"非英语"而系统性变差。
- 什么时候倾向中文:画面里要渲染中文字(招牌、菜单、书法、字幕)时用中文描述并给出中文字符串,通常比用英文转述更准。
- 什么时候倾向英文:指令里夹带大量摄影/美术术语(
three-quarter view、softbox key light、rim light、gouache texture)时,英文术语更稳。 - 别混着乱写:一条指令里主语和对象保持一致的语言即可;给画面文字时引号内的内容必须是原样字符,不翻译。
- 用户用中文提需求时,直接用中文写指令是最自然的选择,不必额外翻译成英文。
5. 可抄模板
5.1 换背景(保留主体)
Replace the background of <image1> with a sunset beach: warm low sun near the horizon, calm water with soft reflections, gentle haze. Keep the subject in <image1> exactly as it is — same pose, same clothing, same hair, same lighting direction on the face.
5.2 换装 / 换材质
Change the jacket on the person in <image1> to a deep emerald green corduroy jacket with the same cut and length as the original. Keep the person's face, hairstyle, pose, background and lighting unchanged, and match the new fabric's sheen and folds to the original light direction.
5.3 多图合成(人 + 服装 + 场景)
Place the person from <image1> into the street scene from <image3>, wearing the outfit from <image2>. Keep <image1>'s face, hair and body proportions unchanged; match the street's ambient lighting and colour temperature on the subject; keep the perspective of <image3> so the person stands naturally on the ground.
5.4 改画面里的文字
In <image1>, replace the text on the storefront sign with "营业中" in the same bold sans-serif style, same size, same perspective and same warm neon glow. Keep the rest of the storefront, the street and the lighting completely unchanged.
5.5 去物并补背景
Remove the parked car from the right side of <image1> and continue the brick wall, pavement and shadow naturally across the area where it was. Keep everything else in the photo unchanged.
5.6 风格迁移
Restyle <image1> as a loose watercolour painting: visible paper tooth, soft wet-on-wet edges, a limited palette of indigo, ochre and warm grey. Keep the composition, subject placement and the person's features recognisable.
5.7 抠主体 → 透明背景
This is an RGBA image with transparency. The product from <image1>, centred with clean crisp edges and a soft contact shadow beneath it. The image has alpha channel and the background is transparent.
5.8 换角度 / 出正视图
Generate a front view of the character in <image1>, matching the same face, hairstyle, outfit and colour palette, in the same art style, on a plain neutral background with even studio lighting.
6. 常见坑
- 编号与传参顺序不一致:
<image2>到底指哪张,取决于reference_images的顺序。写提示词前先把列表定下来。 - 只看
resolution的那句"缩放到 resolution²",以为只影响条件图:它同时决定输出画布。想压条件图成本又不牺牲画布,是做不到的——本工具里两者是同一个 latent。 custom_size打开但megapixels离<image1>太远:官方明确提醒会导致编辑偏移(内容跑到错位置)。要放大就接受这一点并用提示词钉住取景;要精确编辑就保持custom_size关闭。- 指望
negative_prompt起作用:工作流cfg=1,负向无效。要正面描述你想要的画面。 - 一句话里塞五处修改:容易只做其中一两处或整体漂移。拆成多条指令、多跑几次,比一次贪多稳。
- 没说"保持其余不变":编辑幅度小的任务尤其容易被整图重绘;把要保留的要素显式写出来。
- 文字模糊或尺寸不对:文字渲染是强项但不是万能,长段落/小字号仍可能糊;必要时分两步(先出图,再用
edit_image只改字)。 - 分辨率过大:受限于本地条件,
custom_size开到 3MP 这类量级时,图像生成所需时间可能会超过工具等待时间。超时最后会产出文件,但你无法得到。.
Scan to join WeChat group