Office + sporty cast.
This is the final, audio-muxed output—not a raw or intermediate render—and the two source images below are the verified inputs.
Updated October 7, 2026


What changed and what stayed
The two identities and wardrobes come from the source portraits. The staging, gestures, performer positions and timing follow the recognizable reference motion. The final file includes the audio track.
Input requirements
Use one clear, authorized person per image. The service accepts JPG, PNG and WebP up to 10 MB each. Typical generation time is about seven minutes, although queue conditions can vary.
Observed result
The first input maps to the office-look performer and the second maps to the sporty performer. In the 10.06-second final MP4, the pair remain visually distinct through head movement and hand gestures. Clothing color and silhouette remain recognizable, while the template supplies the orange environment, hanging microphone, timing and camera language.
Evidence limits
This case demonstrates one completed run, not a guaranteed result for every portrait. The source sheets are unusually clear, square and evenly lit. Cropped faces, group photos, heavy occlusion or extreme profiles can reduce identity stability. We publish the exact final file so visitors can pause and inspect it rather than relying on selected still frames.