← Back to skills
extension
Category: Content & MediaAPI key requirement unconfirmed

minimax-h3-base

为 dsh 插件的 image_to_video / reference_to_video 工具提供的 MiniMax H3 提示词公共层(i2v 与 r2v 共用):部署事实、时间线写法、镜头与切镜、运镜三要素与词表、说话人 ID 与 <d> 对白/演唱、画面可见文字、overall_soundscape、non_diegetic_music、公共常见坑。需与 minimax-h3-i2v / minimax-h3-r2v 一起加载。

personAuthor: reDDragonlhubModelScope

MiniMax H3 提示词公共层

本技能是 minimax-h3-i2v 与 minimax-h3-r2v 共用的公共层:两个模式共用的规则都放这里,模式专属内容留在各自技能里——minimax-h3-i2v 的指令行与 [Shot 1] 路径写法见其「提示词写法」一节,minimax-h3-r2v 的六段式与四类参考标签同样见其「提示词写法」一节。

第 2–8 节是提示词正文的写法(英文,即要落进 prompt 里的字段与句式),第 1、9、10 节是本地事实与检查清单(中文)。


1. 部署事实

完整系统 = H3-Context-IR(多模态指令理解)+ H3-Base(生成)+ H3-Regenerate-2K。

Context-IR 没有开源,本地这份工作流只包含 H3-Base,所以提示词必须自己写成结构化(写法见 minimax-h3-i2v / minimax-h3-r2v 的「提示词写法」一节)。

没有负向提示词(模型是 CFG 蒸馏权重,工作流里根本没有负向分支),所以别试图用"不要出现 X"来控制画面。

语言偏好:英文>中文。

补充事实:工作流帧率 24fps、输出 32kHz 立体声(minimax_h3_audio_vae_fp32),采样器 res_multistep + simple 调度。(按 i2v 工作流记录;r2v 工作流未逐项核对。)


2. 时间线写法

integrated_multimodal_description is the main body of the rewritten prompt. Every detail should correspond to something visible or audible: visual style, initial composition, subject appearance and position, scene and key props, actions and reactions, shot changes, spoken language, and synchronized diegetic sound.

At the beginning of [Shot 1], state the overall style and initial composition. Common styles include Cinematic, live-action, 2D-animated, 3D CG, claymation, watercolor, and vintage film. For keyframe tasks, derive the style from the reference image; for T2VA, select it from the user's text.

[Shot 1] Live-action, cinematic, a medium-wide shot frames...

3. 镜头与切镜

Do not add a timestamp to the first shot. Use sequential shot numbers for later shots, and begin each one with a strictly increasing cut time that falls within the video duration:

[Shot 2] At 00:03.500, the camera cuts to...

For ordinary cuts, use the camera cuts to, the shot cuts to, the shot transitions to, the shot changes to, or the shot switches to. When explicitly requested by the user, cross-dissolve, fade, or wipe may also be used. A cut should introduce new information about the subject, space, state, viewpoint, or time. If only the distance or a slight angle needs to change, prefer camera motion.

4. 运镜三要素与词表

A complete camera-motion expression has three dimensions: the motion type defines how the camera moves, amplitude defines the range of compositional change, and speed defines the pacing of that change. Add amplitude and speed only when they are meaningful; medium amplitude and normal speed are usually omitted.

| Dimension | Available Expression | Description | |-|-|-| | Motion type | Zoom In / Zoom Out | The focal length changes while the camera body remains stationary | | Motion type | Push In / Pull Out | The camera moves forward / backward | | Motion type | Pan Left / Pan Right | The camera remains in place while the lens pivots horizontally | | Motion type | Truck Left / Truck Right | The camera translates horizontally | | Motion type | Tilt Up / Tilt Down | The camera remains in place while the lens pivots vertically | | Motion type | Pedestal Up / Pedestal Down | The entire camera moves upward / downward | | Motion type | Arc Shot | The camera moves in an arc around the subject | | Motion type | Tracking Shot | The camera follows a moving subject | | Motion type | Static Shot | The camera position and lens remain still | | Motion type | Shake Slightly / Shake Strongly | Slight / strong camera shake | | Motion type | POV | The subject's point of view | | Motion type | Roll Clockwise / Roll Counterclockwise | The camera rolls clockwise / counterclockwise around the lens axis | | Amplitude | with small amplitude | Small-range change | | Amplitude | with large amplitude | Large-range change | | Speed | at slow speed | Slow movement | | Speed | at fast speed | Fast movement |

Camera motion should be written as a natural English action within the shot, rather than stacked as separate labels at the end of a sentence:

The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.
The camera pans right with large amplitude at fast speed, revealing the open doorway.
The camera holds a static shot as the runner exits the frame.

5. 说话人 ID 与 <d> 对白/演唱

Subjects who speak, sing, or produce an off-screen human voice use stable IDs such as (S1) and (S2). When multiple already-numbered speakers speak or sing together, use a compound ID such as (S1,S2). A speaker keeps the same ID across shots; characters who never vocalize receive no speaker ID.

When a speaker first appears, provide enough information from the visual and audio context to establish a stable identity, such as character type, age, gender, whether the person is on-screen, pitch, timbre, speaking rate, or accent. Place the speaker's identifying phrase, ID, action, and delivery outside <d>. Inside <d>, include only the language tag and the actual user-provided spoken content. Preserve every original word and punctuation mark verbatim; do not translate or rewrite them.

The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>

For voiceover, use the exact phrase says in an off-screen voiceover. Immediately after every voiceover <d> block, state that the corresponding on-screen character's lips remain closed:

The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.

When the same line of dialogue or lyrics crosses a cut, use <scenetrans> at the connecting points in both parts and explicitly state that the audio continues across the cut. Use <cutoff> when speech is truncated by the end of the video. Continuity may be expressed with continues seamlessly across the cut, continues uninterrupted into the next shot, carries over from the previous shot, or remains audible across the transition.

6. 画面可见文字

Place any banner, sign, label, subtitle, or neon text that is actually visible on screen in English double quotation marks. Preserve the original text and punctuation verbatim, without translation.

A red neon sign reading "营业中" glows above the doorway.

7. overall_soundscape

Use 1–4 English sentences in one continuous paragraph to summarize the ambient sound, physical action sounds, and non-verbal human sounds across the full video, such as wind, rain, traffic, footsteps, fabric movement, impacts, breathing, laughter, or panting. Dialogue, singing, and diegetic music already belong in the multimodal description and should not be repeated here. Use N/A only when the user explicitly requests complete silence throughout the video.

overall_soundscape: Steady rain taps against the café windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.

8. non_diegetic_music

Use 1–3 English sentences to describe background music that the characters cannot hear and only the audience can hear. Focus on instrumentation, speed, rhythm, and dynamic changes; do not use abstract mood words or explain the emotional function of the score. Singing, instruments, radio, television, or phone music audible to the characters are diegetic events and should appear in the multimodal description. Use N/A when there is no non-diegetic music.

non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.

9. 公共常见坑

  • 只写一句自然语言就发:本地没有 Context-IR,画质/指令跟随会明显下降。先按结构重写再发。
  • 用了否定句:没有负向提示词,写"不要出现文字"没用,要正面描述你想要的画面。
  • [Shot 1] 带了时间码:第一个镜头不带时间码,后续才写 At MM:SS.mmm。
  • 切镜时间超出时长:At 00:08.000 只对 8 秒以上的视频成立;先定时长再写时间码。
  • 对白写进了 overall_soundscape:对白/演唱只能在主体描述里,且只能在 <d> 里给出完整内容。

10. 附:两模式共用的补充条目

  • <d> 是 H3 用于包裹实际说出的内容的特殊 token,格式为 <d>[语言标签] 原话</d>。
  • 语言标签是必要的。对白稳定支持 11 种语言:[Arabic]、[Chinese]、[English]、[French]、[German]、[Italian]、[Japanese]、[Korean]、[Portuguese]、[Russian]、[Spanish]。其他语言也支持,但程度不一。