Rule 1: Put every spoken line inside angle brackets, for example <Welcome to the empty mall.>Rule 2: For longer visual videos, put each numbered scene on a new line, for example Scene 1, Scene 2, and so on. Scene time is estimated from its actions and dialogue; detailed scenes may span multiple continuation clips.
The backend extracts one frame as the character reference and the video's speech as the Seed-VC voice reference. A separately uploaded image or voice sample overrides the matching part of the video.
The generated speech keeps its timing, then Seed-VC changes the speaker identity to match this voice sample.