← Back to skills
extension
Category: Content & MediaAPI key requirement unconfirmed

minimax-h3-i2v

为 dsh 插件的 image_to_video 工具(video_minimax_h3_i2v 模型)提供提示词写法、参数选择与语言偏好。

personAuthor: reDDragonlhubModelScope

MiniMax H3 图生视频(image_to_video)

前置依赖:本技能只写 i2v 专属内容;公共规则见 minimax-h3-base。 用本技能前先加载 minimax-h3-base,其中与 prompt 直接相关的七节是:「时间线写法」「镜头与切镜」「运镜三要素与词表」「说话人 ID 与 <d> 对白/演唱」「画面可见文字」「overall_soundscape」「non_diegetic_music」。 若尚未加载:调用 skill 工具加载 minimax-h3-base。

工具参数只有:prompt、input_image、output_path、last_image、duration、aspect_ratio、megapixels、seed、steps、enable_lightning_lora。 没有负向提示词(模型是 CFG 蒸馏权重,工作流里根本没有负向分支),所以别试图用"不要出现 X"来控制画面。

1. 提示词写法

以下规则适用于首帧(I2VA)、尾帧(L2VA)和首尾帧(FL2VA)模式。各模式特有的指令行和 [Shot 1] 运动路径写法见下面三节(1.1–1.3)。

三个核心字段名与顺序不能改:

integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...

本节其余公共规则见 minimax-h3-base。

1.1 I2VA / 首帧生视频(只传递首帧图片)

指令行

首帧模式的 prompt 第一行必须是以下固定句式,逐字原样写出(I2VA 专属句式):

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

指令行之后空一行,再接三个核心字段(integrated_multimodal_description、overall_soundscape、non_diegetic_music)。

[Shot 1] 的写法:从首帧出发,向前发展

<Picture 1> 是视频 0.00 秒的实际首帧,属于 [Shot 1]。

描述要先确立参考图中的风格、主体、构图和场景锚点,再描述接下来的动作。角色身份、服装、颜色、关键道具、空间关系需保持一致。

推荐结构:

first-frame anchor → action onset → continuous development → result or reaction

1.2 L2VA / 尾帧生视频(只传递尾帧图片)

指令行

尾帧模式的 prompt 第一行必须是以下固定句式(L2VA 专属句式):

How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.

需要替换两处:

  • N:实际最后一个镜头的序号。<Picture 1> 属于最终镜头([Shot N]),不是 [Shot 1]。
  • S.SS:视频的实际时长,格式化为恰好两位小数(如 5.00、4.50)。

指令行之后空一行,再接三个核心字段。

[Shot 1] 的写法:推断前置状态,逐步收敛到尾帧

尾帧模式下,<Picture 1> 是视频最终帧,不属于 [Shot 1],而是属于最后一个镜头。因此 [Shot 1] 不能直接描述尾帧画面,而要从用户意图和尾帧内容推断一个合理的前置状态,然后描述一条明确的动作与过渡路径,让画面在最后一个镜头逐渐收敛并精确落到尾帧的构图、姿态、位置和相机角度上。

推荐结构:

plausible preceding state → explicit action and transition path → gradual convergence in the final shot → last-frame landing

关键约束:不要从最终状态开始写。 如果 [Shot 1] 就直接描述尾帧画面,模型会以为视频从头到尾都是静止的尾帧状态,失去"推断开场"的意义。

1.3 FL2VA / 首尾帧生视频(同时传递首帧和尾帧图片)

指令行

首尾帧模式的 prompt 第一行必须是以下固定句式(FL2VA 专属句式):

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.

需要替换两处:

  • N:实际最后一个镜头的序号。<Picture 2>(尾帧)属于最终镜头([Shot N]),不是 [Shot 1]。
  • S.SS:后端实际采用的有效时长,格式化为恰好两位小数(如 5.00、4.50)。注意:不能直接用用户输入的名义时长,必须用经过帧网格换算后的实际时长。

指令行之后空一行,再接三个核心字段。

[Shot 1] 的写法:描述两端之间的连续路径

<Picture 1> 是视频 0.00 秒的首帧,属于 [Shot 1]。<Picture 2> 是视频最终帧,属于最后一个镜头([Shot N])。

首尾帧模式的核心任务是补全两个端点之间的连续变化,而非把两张图各描述一遍。描述要聚焦于:主体如何运动、姿态如何变化、物体如何被操纵、构图如何演进、场景或光照如何过渡。

推荐结构:

first-frame state → observable intermediate changes → progressively narrowing differences → last-frame state

关键约束

1. FL2VA 一般偏好单镜头。 这样模型才能从首帧连续插值到尾帧(FL2VA generally favors a single shot so the model can interpolate continuously from the first frame to the last frame);只在用户明确指定切镜时才使用多镜头。

2. 最后一个 [Shot N] 必须在视频结尾精确落到尾帧。 无论中间有几个镜头,最终画面必须与 <Picture 2> 的构图、姿态、光照和空间关系一致。

2. 工具参数说明以及使用场景

| 参数 | 工作流默认 | 怎么用 | | --- | --- | --- | | prompt | 必填 | 已在上文详细说明 | | input_image | — | 首帧图绝对路径 | | last_image | 无 | 尾帧图绝对路径 | | duration | 5 秒 | 一般保持默认,必须 > 0 | | aspect_ratio | 1:1 (Square) | 8 个枚举值之一。注意:画面比例不跟随首帧图,是按这个比例重新算画布的 | | megapixels | 0.4 | 与首帧分辨率无关,只决定画布总像素。| | steps | 关 LoRA:20;开 LoRA:8 | 由 enable_lightning_lora 决定数值。所以别在不启用 LoRA 时给 8(会拿到 8 步的完整模型结果,质量塌),也别在启 LoRA 时给 20 | | enable_lightning_lora | false | 打开后走 minimax_h3_fl2v_turbo_8step,速度快但细节/运动幅度会打折。要和 steps=8 配对 | | seed | 随机 | 想复现同一段视频就填同一个值 |

3. 可抄模板

3.1 产品广告(首帧 → 三镜头 + 音频)

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium shot opens exactly on <Picture 1>, the transparent gaming mouse resting on a dark reflective surface in a pitch-black studio void, lit by duotone electric-blue and warm amber rim light. The camera pushes in with small amplitude at slow speed as the internal metallic micro-components brighten. [Shot 2] At 00:02.000, the camera cuts to an extreme macro profile of the ridged scroll wheel; the camera glides along the side at slow speed while a sharp beam of amber light sweeps across the metallic textures. [Shot 3] At 00:03.600, the camera cuts to a low-angle beauty shot; the mouse levitates a few centimetres above the surface and rotates in a slow precise orbit as the duotone light flares along the transparent edges before the frame fades to a silhouette.

overall_soundscape: A deep pulsing sub-bass room tone continues throughout. A sharp tactile mechanical click lands on the first cut, followed by a sweeping glassy whoosh and a rising electronic swell that decays to near-silence.

non_diegetic_music: A sparse analogue synth pulse at a slow tempo, gradually increasing in volume before cutting out on the final fade.

3.2 人物口播(首帧 + 对白 + 画外音)

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium close-up keeps the woman in <Picture 1> in the same seat, preserving her face, hairstyle, jacket and the carriage layout. The camera holds a static shot. She looks up from her phone and the warm, slightly husky woman (S1) says: <d>[English] We arrive in ten minutes.</d> Her lips stop moving as she turns toward the window. [Shot 2] At 00:03.000, the camera cuts to a close-up of her reflection in the rain-streaked glass while her voice carries over from the previous shot.

overall_soundscape: Steady rail clatter and a low ventilation hum continue throughout. Fabric rustles as she shifts in her seat and rain ticks against the window.

non_diegetic_music: Warm muted-piano chords at a slow tempo with a sustained low string underneath, no swell.

3.3 首尾帧过渡(单镜头)

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 5.00-second mark of the target video.

integrated_multimodal_description: [Shot 1] Live-action, cinematic, the shot begins in the exact position and framing established by Picture 1, the cyclist holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the handlebar, raises the umbrella above her shoulder and presses the runner upward until the canopy opens. Water rolls off the expanding fabric, she steps beneath it and rotates the handle into the final angle, settling into the pose, spacing and composition established by Picture 2 at the end of the shot.

overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes.

non_diegetic_music: N/A

4. 常见坑

  • 忘记首帧对齐指令行:首帧任务第一行必须原样写出 <Picture 1> ... fully referenced,后面空一行。
  • 想让画面比例跟随首帧图:做不到,比例由 aspect_ratio 决定;要贴合首帧就手动选一个接近的比例。
  • 一次塞三个场景变化:几秒的视频塞多段剧情容易糊。一个想法一段视频,改动一次只动一个变量。