All Products
Search
Document Center

Alibaba Cloud Model Studio:wan3.0 Video Generation Prompt Guide

Last Updated:Sep 24, 2026

Prompt writing practices for wan3.0 video generation, covering formulas, writing techniques, and tuning examples.

Prompt optimization Skill

The platform provides a wan3.0 prompt optimization Skill to help you tune your prompts.

Download Skill: wan3-pe.zip (unzip after download)

Usage: In the AI chat box, enter /wan3-pe + your prompt content to start debugging the prompt. For example:

/wan3-pe Refine the prompt using the skill. Prompt: a cat running on the grass

Task types

wan3.0 supports the following task types. The inputs and outputs of each task are as follows. For detailed input material combination rules, see Material combinations.

Task typeSupported capabilityReference inputOutput
Text-to-videoText-to-videoPromptVideo with single-shot or multi-shot narrative
Image-to-videoFirst frame to videoPrompt + first frame imageVideo generated with this image as the first frame
First and last frame to videoPrompt + first frame image + last frame imageVideo transitioning from first frame to last frame
Reference-to-videoSubject referencePrompt + reference image or video of a person / object / scene / virtual characterNew video preserving the reference subject's appearance / voice
Motion referencePrompt + reference video containing action / expression / camera movement / VFXNew video replicating the reference motion information
Style referencePrompt + style reference image or videoVideo following the reference style
Audio referencePrompt + music / dialogue / voice timbre reference audioVideo whose audio track aligns with the reference audio
Clay render reference / renderPrompt + clay render videoPhotorealistic re-rendered video
Multi-grid shot / storyboard referencePrompt + multi-grid shot imageVideo advancing along the shot storyline
Keyframe referencePrompt + multiple keyframe imagesVideo aligned with the keyframe order
File / webpage referencePrompt + document or webpage linkVideo generated based on the document / webpage content
Video editingVideo instruction editingPrompt + source videoVideo after add/remove/modify of elements per instruction
Video reference-image editingPrompt + source video + reference imageVideo after replacing / modifying the source video with reference image content
Video audio editingPrompt + source videoVideo after adjusting vocals / music / sound effects
Video extensionPrompt + source videoVideo continued before / after / bidirectionally from the source video

Prompt element writing

A single sentence is enough for wan3.0 to produce a video, but the more complete and precise the description, the closer the result is to what you expect. For full-modality scenarios, we recommend the following complete prompt formula — write down whatever you want to control, and simply skip any section you do not need.

Prompt = [Overall description]
       + [Reference material citation: Image N / Video N / Audio N]
       + [Shot N (start-end seconds): Subject + Scene + Motion + Aesthetic control]
       + [Dialogue: xx says: "xxxx"]
       + [Sound effect / BGM]
       + [Style / Mood]
       + [Negative prompt list]

Example

A writing example that combines reference-to-video + multi-shot + dialogue + voice timbre reference + style control. The prompt is strung together into one paragraph following the formula:

Eastern epic xianxia, majestic and imposing. A general in heavy armor holds a long spear, riding alone on a golden qilin, guarding the grand pass between two mountains as an enemy army surges like a black tide in the distance. 15 seconds, pacing shifts from steady restraint to explosive release. The general's appearance references Image 1, facial expression references Image 2, mount references Image 3, voice timbre references Audio 1.
Shot 1 (00:00-00:03): Extreme wide establishing shot, 24mm wide-angle, extreme low angle, camera static. A towering stone archway stands between two mountains; on each side stands a weathered giant stone statue of a warrior. The general rides the golden qilin alone before the gate, tiny yet unyielding. In the distance, the enemy black tide blankets the wilderness and slowly advances; dark clouds press down, fierce winds whip up sky-filling dust, tattered war banners snap in the wind.
Shot 2 (00:03-00:06): Close-up, 85mm telephoto, shallow depth of field. Half of the general's face, blood and dust on the temple, resolute eyes; cut to the hand gripping the spear, knuckles white. The qilin growls low and exhales white breath, golden scales shimmer with flowing reflections. The general speaks low: "This pass, I have guarded for ten years."
Shot 3 (00:06-00:10): Medium close-up, slightly low angle, camera pushes in very slowly. The general raises his eyes, killing intent surging within, and says in a deep voice: "Today, not one of you gets through." As the words fall, the qilin rears its front hooves high and slams them to the ground.
Shot 4 (00:10-00:15): The camera orbits 120 degrees around the general and descends to an extreme low angle, transitioning to an extreme wide shot. A three-headed, six-armed golden dharma form rises behind him, taller than the archway, moving in sync with the general; where the spear strikes the ground, radial cracks explode outward, a ring-shaped dust shockwave expands, the enemy black tide halts as one, and the silhouettes of the general and qilin stand in golden light.
Style and mood: Eastern epic, IMAX-scale spectacle, anamorphic widescreen, cinematic, film grain, high dynamic range; tragic and unyielding, the overwhelming presence of one man holding the pass.
Negative prompt list: No modern elements, no subtitles or watermarks, no extra characters, avoid facial distortion and continuity errors, avoid blur and low resolution.
Input materialOutput video

Image 1

Image 1

Image 2

Image 2

Image 3

Image 3

Audio 1

Reference material citation

Materials are numbered by upload order, with images, videos, and audio counted separately: the first image is Image 1, the second is Image 2; videos and audio follow the same pattern as Video N and Audio N. Therefore, Image 1 and Video 1 can exist at the same time.

  • In the prompt, write Image N, Video N, or Audio N directly to cite the corresponding material — the numbering must match the upload order, otherwise the wrong material will be cited.
  • When there is only one material, you can abbreviate it as reference image or reference video; with two or more, you must specify the number.
  • The same material can be cited in multiple places — cite it separately for appearance, motion, voice timbre, scene, and other uses.
What you want to achieveReference writing
Single materialReference image, generate a video of it running; Reference the camera movement of the video, reshoot a segment
Lock appearance onlyThe child in Image 1 stands in the entryway of the modern apartment in Image 2
Multi-material mixThe figure in Image 1 holds Image 2 in hand and walks into the scene of Video 1
Voice cloningVoice timbre references Audio 1; Speak in the timbre of Audio 1: "xxxx"
Borrow motion / camera movement onlyReference the camera movement of the video, change the scene to a cup of coffee on a rooftop cafe table
Borrow action onlyReplace the skateboarding man in Video 1 with the woman in Image 1, keeping the skateboarding action, trajectory, and scene unchanged
Borrow VFX onlyThe character in Image 1 follows the energy VFX of Video 1, generating the same flame-burning VFX and action
Borrow style onlyStyle references Image 2; Only reference the color tone and lighting of Image 2, not the characters
Document / webpage referenceBased on my proposal document, generate an ad film for a brand; Based on Wang Wei's encyclopedia entry, generate a first-grade science animation

Shots and timestamps

Multi-shot videos use "number + timestamp + shot content" to form a coherent narrative, keeping subject, scene, and atmosphere consistent across shots.

Prompt = Overall description + Shot number + Timestamp + Shot content

  • Overall description: One sentence stating the theme, perspective, narrative style, and overall mood.
  • Shot number: Number each shot to define the order.
  • Timestamp: Mark the start and end time right after the number; each shot connects end-to-end with no gaps or overlaps, 2–5 seconds per segment.
  • Shot content: Describe subject action, dialogue, expression, and scene details, same as a single-shot write-up; avoid high-frequency actions the model cannot control (such as "shaking head 3 times per second").

The number and timestamp format is flexible; [start-end] and (start-end) are interchangeable — pick one and use it consistently throughout:

WritingExample
Number + timestampShot 1 [0-3s] xxxx, Segment 1 [0-5s]: xxxx, Shot 1 (00:00-00:03): xxxx
Number only, no timeShot 1: wide shot xxxx; Shot 2: close-up xxxx
Timestamp only, no number[00:00-00:03] xxxx

Full example:

This story is told from a third-person perspective, a short drama about letting go and rediscovering hope.
Shot 1 [0-3s] A boy sits alone in a corner of the playground, looking down at the letter in his hand, then sighs softly, eyes lost.
Shot 2 [4-6s] Hard cut, fixed camera, focused on the boy's eyes, tears glistening, filled with loss and helplessness.
Shot 3 [7-10s] Hard cut, scene shifts to a simple classroom; a girl with a gentle, firm expression walks to the boy and comforts him.
What you want to achieveReference writing
Single shot, no splitWrite Generate single shot / One continuous shot / Generate single shot. on the first line
Camera completely stillFixed shot, camera static, position unchanged
Specify camera movementUse plain language, e.g. push in, pull out, orbit, handheld follow
Transition between shotsWrite hard cut or dissolve at the end of a segment, or start the next segment with seamlessly continue from the last frame of the previous segment
Emphasize pacingAdd to the overall description pacing shifts from steady restraint to explosive release / eight shots total, brisk buildup, sudden conflict, short climax, warm ending
Emphasize focal length / depth of field85mm telephoto, shallow depth of field / 24mm wide-angle, deep focus
High-speed photographyHigh-speed photography (1000fps slow motion)

Dialogue

Dialogue is written as "character + speaking verb + colon + quoted content", e.g. xx says: "content", she whispers: "content".

  • Multi-person dialogue: Write line by line and identify the speaker, e.g. Image 1 says: "xxxx"; Image 2 replies: "xxxx".
  • Voice cloning: Append Voice timbre reference Audio N — write it once and it applies to all of that character's dialogue.
  • Lip sync: Append Lip sync or Lip sync..
  • Voiceover, monologue: Same format, e.g. Voiceover: "xxxx"; Female voice chanting: "xxxx".
  • No dialogue: Explicitly write No dialogue or No dialogue., otherwise the model decides on its own whether to include dialogue.

Sound effects / BGM

Ambient sound auto-matching: As long as the action, material, and weather are clearly described in the frame (war banners snapping, white breath from nostrils, splashing through puddles), the model will automatically fill in the corresponding sound effects.

Specifying sound effects as standalone sentences, optionally with timestamps:

Gunshot sound effect + fox-fire particle VFX
Whoosh on color-block slide-in, light keyboard click on number jumps
From the 8th second, female voice chanting: "xxxx"

BGM three-level control:

  • Conservative / open: Write nothing and let the model choose.
  • Explicitly specify: BGM drops to the low register, cello enters very softly and slowly, Epic symphony + electronic sound effects mixed, rhythm in sync with shot cuts.
  • None at all: No BGM, generate only ambient and action sounds or No background music..

Style / Mood

The three-part set of art style + color tone + mood can be placed at the beginning or end of the prompt; you can expand into a detailed description or just give keywords.

SectionKeyword reference
Art style35mm cinematic film / Hong Kong realism / line illustration / wasteland style / Miyazaki animation / Chinese-style mineral-color painting / clay animation / motion graphics / Apple Keynote minimalist business style
Color toneHigh-saturation warm yellow / Cyan-blue × gilded gold contrast / Low-key high-contrast lighting / Desaturated ambient cyan-gray against high-saturation gold
MoodTragic and unyielding / Light and healing / Oppressive and tense / Relaxed and confident / Mysterious and magical

Negative prompt list

Write only what you do not want to appear; leave it empty if there is none. Do not pad the list or repeat content already stated in the positive prompt.

Common negativeExample
Characters / frameNo face distortion, No face swapping, No extra fingers, No twisted limbs, No clipping artifacts, No cutout traces, No low clarity
Text / watermarkNo subtitles, No watermark, No complex text, No extra brand names
StyleNo childish cartoon, No photorealistic live-action, No game-CG look, No music-video showboating, No cheap cyberpunk, No modern high-definition digital look
SoundNo background music, No voiceover, No dialogue, No background music dominating the mix
ActionNo simple tearing, No characters running illogically, No exaggerated limb movements, No exaggerated blinking, No slow motion
Camera movementNo large camera movements, No exaggerated lens rotation, No sustained camera shake, No over-stabilization, No frequent shot cuts
SpaceNo spatial chaos, No vehicle clipping, No axis errors, No multiple people in frame
ConsistencyNo character position changes, No costume changes, No scale drift (common in editing / reference scenarios)

Prompt scenario examples

Multi-subject consistency

Pass in multiple people, prop, and scene materials at once, and lock them individually with Image N / Video N / Audio N. Clearly separate the purpose of each material in the prompt to prevent features from bleeding across subjects.

Reference materialPromptOutput video

Image 1

Image 1: little girl

Image 2

Image 2: lollipop

Image 3

Image 3: punk

Image 4

Image 4: chameleon man

Image 5

Image 5: qipao girl

Image 6

Image 6: tiger man

Image 7

Image 7: Chinatown alley
Urban fantasy kung-fu action film scene, a Chinatown alley at dusk. Hong Kong-style realistic cinema look, high-saturation warm yellow lantern light accented with red and green neon signs, high contrast, slight film grain, authentic photographic texture; the mood shifts from light and brisk to tense, then settles back to warm. Twenty-five seconds in total, eight shots, pacing of light buildup, sudden conflict, brief burst, warm ending. The performance and fighting must feel real and weighty, like a big-budget Hong Kong action scene; no video-game CG feel, no MTV-style showboating, no heroic slow-motion posing entrance.
Subject locking: little girl Image 6, lollipop Image 7, punk Image 4, chameleon man Image 5, qipao girl Image 2, tiger man Image 3, Chinatown alley Image 1.
The scene is in the Chinatown alley Image 1: a narrow lane, both sides hung with vertical Chinese signboards and iron window grilles, overhead strung with red lanterns and clotheslines, paper boxes and plastic crates piled at the wall base, the ground wet with puddles reflecting light. The little girl is about six, pink dress, two small braids, holding up a colorful lollipop. The punk is a tall thin youth, green hair, floral shirt, red shorts, walking with a swagger. The qipao girl is in her early twenties, deep blue dark-patterned qipao, upright posture and calm expression. The alley deepens into darkness toward the back, while the alley mouth lets in orange-gold dusk skylight.
Shot 01, wide shot establishing tracking. 24mm wide-angle low angle near the ground, camera slowly dollies forward behind the little girl. The girl walks while licking the lollipop, steps light and brisk, red lanterns sway gently overhead in the wind, clothesline garments flutter. The base layer is distant main-street market noise, wok sizzling, Cantonese voices and bicycle bells, muffled by the alley.
Shot 02, medium shot, alley depth reveal. 50mm static camera, framing with a slight Dutch angle to create unease. The punk steps out of shadow, first a backlit silhouette, then two steps closer to reveal his face, a nasty smile with eyes locked on the lollipop. Punk: "The candy, give it." Market noise dips, leaving only his footsteps and the shuffle of slippers on the ground.
Shot 03, close-up reaction shot. 85mm telephoto shallow depth of field, background blurred into lantern bokeh. The little girl stops, pulls the lollipop toward her chest, steps back half a pace, expression of startle and wariness, not terrified screaming.
Shot 04, qipao girl enters. Frame-within-frame composition, camera shoots from a side arched doorway into the alley. The qipao girl walks in from the side of frame, unhurried, the qipao hem swaying lightly, stops half a step in front of the little girl, shielding the child behind her. Qipao girl: "Don't bully the kid." The action should feel like an everyday casual interruption, no slow-motion posing, no striking a stance.
Shot 05, a kick to the wall. Handheld tracking with a fast whip pan allowing brief motion blur. The girl lands a front kick squarely on the punk's chest, the qipao slit lifted by the leg wind. The camera whips to follow the direction the punk is kicked, he slams hard into the brick wall, wall plaster flakes off, paper boxes crash down, overhead lanterns sway violently. The little girl startles back, the lollipop slips from her hand, falls and rolls twice.
Shot 06, transformation and tongue strike. Close-up of the punk's face, skin ripples upward from the neck, color shifting from skin tone to ink-green and blue-gray, eyes rotating separately, pupils splitting vertically, becoming chameleon man Image 5. The camera pushes in extremely fast. He opens his mouth and a very long tongue snaps out straight at the qipao girl. The transformation must be a clearly visible layer-by-layer skin-color ripple, not a blurry blob of light.
Shot 07, transformation and grappling. The girl turns sideways to dodge the tongue, the tongue tip grazes the wall beside her ear, brick dust scattering. She lets out a low growl, orange-black tiger stripes surfacing along her neck and the backs of her hands, nails becoming sharp claws, irises turning amber, becoming tiger man Image 3. The camera slowly orbits the two maintaining handheld breathing shake; the two trade rapid blows in the narrow alley, claw shadows and tongue whips crossing under the lantern light, moves brief and solid. The clothesline snaps, garments fall, one lantern is smashed, fire scattering on the ground.
Shot 08, a palm strike sends him flying and the ending. Extreme low angle upward shot. The tiger man strikes with one palm on the chameleon man's chest, the palm wind kicking up a visible shockwave and dust, the chameleon man flies backward into the shadow at the alley's end and is gone. The tiger man settles into stance, the tiger stripes and claws slowly fade, returning to the qipao girl's appearance. She turns and crouches, picks up the lollipop from the ground, blows on it and hands it back to the little girl. The little girl smiles, takes the candy and holds her hand. The camera slowly pulls back and rises slightly, the two figures walk down the alley toward the warm light at the alley mouth, red lanterns swaying quietly overhead once more.
Physical effects throughout must be clearly visible: dress and qipao hems swaying with motion, lollipop wrapper reflection, red lantern swaying and shattering, clothesline garments falling, steam and floating dust in the alley, wall plaster flaking, paper boxes collapsing, flyer scraps blown by the shockwave, chameleon man skin-color ripple and wet tongue-surface reflection, tiger man stripes surfacing and neck-side fur bristling, palm-wind shockwave and ground dust rising, neon reflections on the wet ground. Lighting uses overhead red lanterns and side signboards as practical sources, alley-mouth dusk skylight as backlight rim, red and green neon light spots dotting the wet walls and puddles, hard light mixed with warm light, high saturation and high contrast. In the fight passage a lantern shattering may cause a brief light flicker, but no strobing.
No injury, bleeding or excessive terrified crying from the little girl; no gory brutal fighting detail; no flickering, shaking, frame-skipping or self-inserted cut points; no distorted or drifting faces; no transformation collapsing into a blurry light blob; no chameleon man tongue breaking, sticking or deforming out of control; no tiger man turning into a full four-legged beast, must remain humanoid with tiger features; no extra fingers or limb misalignment; no text or watermarks.

Image 1

Image 1: red fox dual-pistol gunner · Fei Shou

Image 2

Image 2: human alchemist · Zi Luo

Image 3

Image 3: tiger-striped beast-man mage · Ting Yue
[Video type] Game user-acquisition character PV showcase video
[Video duration] 15 seconds
[Reference material] Three operator character illustration images: red fox dual-pistol gunner Fei Shou Image 1, human alchemist Zi Luo Image 2, tiger-striped beast-man mage Ting Yue Image 3
[Shot structure]
Segment 1 (0-3s): Three characters appear in sequence in close-up shots, emerging from shadow, signature elemental particles swirling around them (Fei Shou Image 1 fox-fire phantom flames, Ting Yue Image 3 thunder-carved spellcraft, Zi Luo Image 2 alchemy essence), the camera slowly pushes in to focus on the face, expressions revealing character traits (Fei Shou Image 1 quick glancing back, Ting Yue Image 3 steady gaze, Zi Luo Image 2 knowing smile)
Segment 2 (3-10s): Skill showcase dynamic shots, rapid cut transitions
- Fei Shou Image 1: side-crouched low, both flintlock long guns firing in succession, mint-gold fox fire bursts from the muzzles and condenses into three fox-shadow clones weaving through in assault, gunshot sound effect + fox-fire particle effect
- Ting Yue Image 3: raises a magic staff at a slant, summoning a golden eagle to spread wings and soar, dark-green bronze thunder erupts from the staff tip sweeping across to awe enemies, thunder-crash sound effect + lightning particle effect
- Zi Luo Image 2: raises a brass balance scale in one hand, violet essence fluid streams from fingertips, condensing into flowing replicas that orbit around, essence surging sound effect + violet translucent particle effect
The camera follows elemental trajectories with pans, character action poses shown dynamically, expressions focused in combat state
Segment 3 (10-15s): Three stand shoulder to shoulder in a group shot, slowly pulling from wide shot to medium shot, the three signature elements converge and blend behind them (fox fire + thunder eagle + essence blending into a particle cloud), color tone gradating (mint gold + dark green bronze + violet), each character showing their trait expression, the three names flashing in sequence
[Expression and dialogue]
- Fei Shou Image 1: quick glancing back, slightly sharp, eyes playful yet assured, dialogue "Under the trigger, no one escapes" with lip sync
- Ting Yue Image 3: steady gaze, slightly commanding, eyes sharp as an eagle, dialogue "Where the falcon reaches, there is thunder" with lip sync
- Zi Luo Image 2: knowing smile, slightly mysterious, eyes warm and calm, dialogue "All things can be equivalent, including fate" with lip sync
[Visual style] Dark tone + high contrast, Fei Shou mint gold, Ting Yue dark green bronze, Zi Luo violet; fox-fire particles, thunder-eagle particles, essence particles converging and blending; game logo emerging, character names flashing, skill names labeled; shadow background → combat scene → high-altitude thundercloud background gradation.

Image 1

Image 1: girl A face reference

Image 2

Image 2: girl B face reference

Image 3

Image 3: black leather suit

Image 4

Image 4: champagne gold silk dress
Cinematic dual-heroine fashion film, extremely fast pace, pure hard cut editing logic, shot on Arri Alexa 65 and Panavision anamorphic widescreen lenses, Kodak Vision3 500T film texture, 8K resolution, photorealistic, no text. Facial features fully locked to girl A Image 1 and girl B Image 2.
[0-3s: high-speed entrance and contrast] (hard cut) Pure white minimalist space, cut by strong side light into sharp shadows. Left screen / left side: girl A Image 1 wearing the black deconstructionist leather suit of Image 3, cold beauty under cold light, fast turn of the head staring into the lens. Right screen / right side: girl B Image 2 wearing the champagne gold flowing silk dress of Image 4, languid eyes under warm light, fast turn of the head staring into the lens. Authentic film grain.
[3-6s: extreme material close-up] (hard cut) Extreme macro close-up. Girl A referencing garment Image 3: cold light on the sharp folds and metal hardware of the black leather, a cold premium luster. Girl B referencing garment Image 4: warm light tracing her collarbone line, silk fabric flowing like water ripples, skin showing a hyper-realistic translucent dewy texture, fine peach fuzz visible.
[6-9s: tension crossing (high-speed photography)] (hard cut) The two walk toward each other in the pure white space, passing shoulder to shoulder. High-speed photography (1000fps slow motion). The hardness of the black leather and the softness of the champagne silk violently collide in the crossing instant, hems and hair flying in the air. In the 0.1 seconds of crossing, their eyes meet sharply.
[9-12s: back-to-back pressure] (hard cut) The two stand back to back. Strong side backlight (Rim light) simultaneously outlines both faces and a golden rim on their hair. They both turn their heads fast, staring into the lens. Expressions intensely pressuring, professional, radiating the confidence of top supermodels.
[12-15s: ease and freeze frame] (hard cut) The two stand shoulder to shoulder, slightly turning heads to look at each other, showing a highly infectious, relaxed and confident smile. The camera slowly pulls back (Slow zoom out), revealing the two's perfect proportions in a grand light-and-shadow space, the image fading in premium glow.
Camera movement: fast handheld, extreme push-in, high-speed photography slow motion, pure hard cuts, no flashy visual effects. Color: high-contrast black-white-gray base, accented with cold metallic luster and warm champagne gold, cinematic color grading.

Image 1

Image 1: urban space and neon atmosphere reference

Image 2

Image 2: male character image reference

Image 3

Image 3: female character image reference
Generate an approximately 30-second, 16:9 landscape, high-aesthetic urban MG chase short film. Image 1 serves as the reference for urban space, neon atmosphere, lane structure and overall visual texture; Image 2 as the sole image reference for the male character; Image 3 as the sole image reference for the female character. Keep both characters' faces, hairstyles, body types, garment structure, color scheme and recognizability stable and consistent.
[Overall style] High-end motion graphics dynamic illustration style, fusing flat MG graphic language, refined 2.5D spatial layering and restrained 3D volumetric feel. Not realistic live-action cinema, not childish cartoon, not grimy heavy-industrial cyberpunk. The tone should be showy, cool, beautiful, tense and eye-catching, like an action-movie-grade urban chase promo. The camera makes heavy use of handheld tracking, fast pans, snap zoom, low ground-hugging angles, forward retreating, whip pan, occasional drone wide shots, to build a strong sense of chase pressure.
[Character action logic] The woman's advantage is being light, fast, agile and good at using the environment and changing lines; the man's advantage is power, anticipation, pressure and skill at interception and close control. Both are martial-arts experts, so the contest is not chaotic grappling but a "grab, block, deflect, switch hands, borrow force, turn, evade, counter" game revolving around the light orb.
[Light orb] The light orb fits in one hand, white core with a soft violet-pink-blue halo, the single object of contest in the whole film. The orb must realistically light the hand, the edge of the face, clothing and the surrounding environment. There is always only one orb, stable in size, not duplicated, not vanished, not deformed; when moving it leaves only a short clean light trail.
[0.0-4.0s | City traversal and woman's entrance] The camera starts from a high-altitude city night skyline, sweeping past skyscrapers, neon signs and road lights, then dives along the main road into the district, passes over traffic and rapidly drops to street level. When the camera stops abruptly, the woman dashes in from the side, clutching the light orb. First a quick set of close-ups: footsteps, the hand gripping the orb, a glance back, the upper body leaning into the sprint.
[4.0-8.0s | Man intercepts, first close contest] As the woman bursts into a narrow corner passage, the man cuts in ahead from another route. The man first reaches for the woman's orb-holding wrist, clearly targeting only the orb. The woman immediately tucks her arm to shield the orb while borrowing the grabbing force to turn, completing the first deflection. The man instantly switches to the other hand to cut off the orb's path and uses his shoulder to seal her way. The woman drops her center of gravity, slides low under the man's arm, and finally shakes him with a quick spin, continuing to sprint out.
[8.0-13.0s | Into the crowd, woman turns the tables] The woman dashes into a dense crowd, actively choosing a complex environment. She first slips sideways between two pedestrians, then uses roadside stalls and railings for quick changes of direction. The man briefly catches the back hem of the woman's jacket in the crowd; the woman fluidly half-turns, redirecting the jacket's pull to the side, while pressing the man's forearm away with a backhand elbow, then slipping past his side.
[13.0-16.5s | Into traffic, mounting a bike] The woman grabs the handlebar of a roadside bicycle and, using her forward momentum, mounts it directly. The man likewise snatches another roadside bicycle to give chase.
[16.5-24.0s | Traffic chase and third contest] The woman's cycling line is more agile; she weaves along the edges of cars, leans into turns, cuts diagonally. The man first draws alongside in the cycling section and reaches for the orb; the woman deliberately keeps the orb on the outside for half a beat, then once the man's weight shifts she suddenly leans and changes line, making him grab empty air. On the second approach, the man uses a decelerating car to block the woman's outward escape and nearly catches her orb-holding wrist. This time the woman doesn't resist head-on but quickly switches the orb to the other hand, slipping out of the man's reach.
[24.0-27.5s | Final pressure and the woman's escape] The intersection ahead has more complex traffic. The woman spots a large bus passing crosswise and a row of hard-stopping vehicles, and instantly judges this as her chance to escape. The woman suddenly leans into a turn, cutting outside a hard-braking car and threading through the narrow gap between the car's front and the bus. The man reacts a beat slower, forced by the angle to slow and detour.
[27.5-30.0s | Ending] Once the bus has passed, the man bursts out only to find the woman has opened the gap. At the distant neon corner only one fast-moving light-orb dot remains. Finally the camera slowly rises and pulls back, showing the bustling night city, continuous traffic and the distant fading light dot; the woman has escaped successfully.
[Negative limits] No smooth ad-film camera, no constant side tracking, no slow motion, no simple grappling-style orb snatching, no characters running illogically, no childish cartoon, no realistic live-action style, no war/police/weapons, no characters changing faces, no orb vanishing/duplicating/deforming, no cheap cyberpunk.

Image 1

Image 1: product front reference

Image 2

Image 2: product back reference

Shot 1: wide shot, the Cuiyan Society black-truffle sea-salt mixed-nut crisp packaging of Image 1 is displayed centered in a kitchen wooden-table scene, with a few almonds, cashews, walnut halves and black-truffle slices scattered beside it as atmosphere props; the camera slowly pushes in closer. Shot 2: close-up, the front of the packaging fills the frame, clearly showing the top gold-stamped wheat-ear nut circle badge logo, the "Cuiyan Society" gold-stamped calligraphy wordmark, the "Black Truffle Sea Salt Mixed Nut Crisp" product name, three capsule labels "Selected 6 Imported Nuts", "0 Trans Fatty Acids", "Non-fried · Low-temp Light Bake" and the deep-forest-green base with gold-stamped color scheme, text sharp and legible. Shot 3: a hand picks up the packaging from above and flips it to the back of Image 2, showing the densely packed ingredient list, nutrition facts table, manufacturer information and bottom barcode, off-white small text clearly readable on the dark green base, then flips back to the front. Throughout, the packaging's gold-stamped logo, text, deep-forest-green color scheme and printed pattern stay highly consistent without deformation. E-commerce product showcase style, warm-tone lifestyle feel.

Storyboard / multi-grid

Upload a nine-grid / stick-figure storyboard image as a plot outline; the prompt fills in the shot, action, sound effect and mood for each panel in order, and the model generates the video panel by panel.

  • Declare usage up front: At the start of the prompt, state that the storyboard is only for shot guidance, e.g. Use the storyboard only as shot guidance; do not generate video from the storyboard as a single image, to prevent the model from treating the whole grid image as a single-frame input.
  • Restate panel by panel: Describe each panel's shot size, camera movement and subject position in left-to-right, top-to-bottom order; environmental, lighting, particle and other details not shown in the grid are supplied by the prompt.
  • Segments need not map one-to-one to panels: The storyboard is only a plot outline; the number of segments can be flexible.
  • End with a unified declaration of art style, render texture and aspect ratio, and attach a negative prompt list to exclude unwanted styles, e.g. No 2D hand-drawn, no Japanese anime, no live-action footage.
Reference storyboardPromptOutput video
Mushroom forest nine-grid storyboard
0s-2s: Crane overhead shot slowly descending, wide shot of a giant mushroom forest — red polka-dot mushroom caps tower like trees, the ground carpeted with glowing moss and dewdrops; a human girl shrunk to thumb size (about 8, brown twin tails, wearing a green hooded cloak) is climbing onto the cap of a mushroom, her companion — an equally mini white tiger cub (blue eyes, plump Pixar-style look) — curiously sniffs golden glowing moss at the mushroom's base; warm and healing atmosphere.
2s-4.5s: Medium tracking shot, the girl stands atop the mushroom cap gazing afar and notices a patch of purple-fluorescent flowers glinting beautifully in the distance; she excitedly waves to the white tiger, the tiger leaps onto her shoulder, and the two slide down the curved mushroom stem and dash toward the flower patch, the camera with slight motion blur.
4.5s-7s: Low-angle upward push-in, the two reach right in front of the purple flower patch; the girl reaches out to touch a petal — the petals snap open revealing a mouthful of fangs and a magenta interior, mutating into a giant Venus flytrap that lunges to bite her; the girl shrieks and leans back to dodge, the white tiger arches its back and roars.
7s-9.5s: Aerial shot rapidly rising, the flytrap's roots burst from the ground and turn into countless purple-fluorescent spiked thorn vines spreading rapidly across the ground chasing the girl and tiger; where the vines pass the moss withers black.
9.5s-12s: Steadicam lateral tracking, the white tiger carries the girl weaving and leaping between collapsed giant mushrooms; vines wrap around mushroom stems and yank them down with crashing debris; the two dodge through narrow gaps while vines close in from all sides.
12s-14s: Close-up rapid push-in, the girl looks down and spots a dark small cave entrance in a mossy rock crevice at her feet; she grabs the tiger and leaps in, the entrance instantly sealed by purple vines.
14s-15s: Static long shot inside the cave, pitch black, only the white tiger's blue eyes glowing softly in the dark; the girl clutches the tiger curled deep in the cave breathing quietly, the sound of vines rubbing outside gradually fades, the frame holds.
No subtitles or character dialogue throughout, only environmental sound effects and action sound effects. The score follows the plot: the opening is light, warm and soothing, the mutation stage abruptly turns to tense uneasy strings, the escape stage is tense urgent percussion, the ending returns to quiet with only breathing and the fading rub of distant vines. Pixar 3D animated film texture, all characters and scenes are 3D CG-rendered — characters have soft rounded 3D volume, stylized subsurface-scattering skin, large expressive Pixar-style eyes, hair as 3D models, clothing fabric with thickness and natural drape; the miniature mushroom-forest scene is rich in moss dewdrop spore particles, strong depth-of-field blur highlighting the macro feel, volumetric light piercing the mushroom caps casting light pools. 8K ultra-high-definition image quality, 16:9 landscape, crane + aerial + Steadicam + close-up rapid push-in composite camera work. No live-action footage, no realistic skin texture, no 2D hand-drawn flat fill, no Japanese anime style, no readable text subtitles or watermarks of any kind.

Image 1

Lu Mei washing dishes dancing · sketch storyboard

Image 2

Lu Mei · three-view character sheet

Use Image 1 as an advertising storyboard to guide the shoot; do not generate video from the Image 1 storyboard as a single image. Use Image 2 as the character, and shoot the full process of "Lu Mei," a yellow cartoon little hippo wearing a lace bonnet, overalls and a small apron, dancing while washing dishes in the kitchen (her performance is very cheerful and soft-cute, dancing with hip-twirls and spins while washing dishes, lifting the dish brush as a microphone to sing, limb movements exaggerated and lively, referencing Disney character-animation performance style); motion and camera movement natural and fluid.

[Visual texture] Cinematic shots, 3D render texture, C4D render texture, Octane renderer texture, Pixar / Ghibli-style soft-cute cartoon, soft cinematic lighting, cinematic color grade, film grain, depth of field;

[Music & sound effects] [Playful child-voice humming] Simultaneously generate lighthearted cheerful background music and sound effects, matching the camera and character performance: a breezy celesta and marimba lead melody, the splashing water of dishwashing, the soft pop of bubbles, the ding of notes lifted by the spin, and a final soft-cute "phew~".

Clay render reference / rendering

Upload a pure-white clay render video as the sole blueprint; replace only the material, color and texture with real-world texture, keeping everything else in the frame unchanged frame by frame.

  • State locked items explicitly: Declare "change nothing except material and color," and list camera trajectory, cut points, object positions, character actions and the like as "absolutely unchanged," repeatedly emphasizing "no additions, no deletions, no moves" and "not one frame added or removed, not one second out of place."
  • Emphasize realism: Use optical details like realistic glass refraction, bark transmittance and leaf veins, skin subsurface scattering to give the generated image a more authentic texture.
Input clay render videoPromptOutput video
Using Video 1 (my uploaded pure-white clay render animation) as the sole blueprint, do one realistic re-render: replace every object and character still in pure-white matte raw form with real-world material and color, rendering into a realistic cinematic live-action texture; change nothing except material and color.
[Absolutely unchanged · highest priority] Camera motion trajectory, the speed and direction of every push/pull/pan/tilt/orbit, the timing of shot-size switches, cut points, aspect ratio, total duration, composition and perspective — all frame-for-frame identical to the reference video, not one frame added or removed, not one second out of place. The position, orientation, size ratio, quantity and mutual occlusion of every object in the scene are copied exactly; no additions, no deletions, no moves. The characters' bone structure outline, hairstyle volume, body proportions, garment tailoring and all actions, walking rhythm, the timing of crouching, reaching and pushing doors must be completely identical to the clay video; no action changes, no swapping of characters. Sunlight direction, shadow angle and length, and the light-dark distribution of every frame are copied from the original. The only thing allowed to change is material, color, texture, and the real optical behavior that results.
[Realistic conversion] After realistic conversion, the clay model's cartoon-rounded forms retain their volume outline and proportions, but every surface is replaced with real material: plastered walls are real cement troweled texture with rain streaks, efflorescence white bloom and moss at the wall base; the tile roof is fired clay tile with color variation, chipped edges and dust accumulation, dry grass and moss in the tile gaps; doors and windows are old real wood with clear grain, worn and cracked paint, glass with real refraction, fingerprints and slight smudges reflecting sky and tree shadows; steps are rounded granite with polished trodden edges; wooden buckets and crates are old pine with splinters and nail holes, iron hoops and nail heads with real reddish-brown rust layers and rust runs; the tricycle is chipped blue iron sheet showing primer, galvanized parts with white rust, rubber tires with worn treads and caked mud; the well is unevenly jointed bluestone, hemp rope with fraying and broken strands, the well platform stained dark by water soaking; tree bark has real cracked furrows and lichen patches, leaves are real layered foliage with bug holes and transmittant veins; ground flagstones have real cracks, color spots and soil in the gaps, grass is real blades with dried-yellow tips; brick walls have real fired color variation and hollow mortar joints.
[Character realistic conversion] The little girl becomes a real child: a real East Asian child's face and skin, skin with real pores, peach fuzz and subsurface scattering, natural flush on cheeks and ear rims, eyes with dark-brown irises showing real texture and wet reflection, lashes and brows as individually distinguishable natural dark brown, hair as real strands in bundles with flyaways and frizz, hair tips catching a light-brown edge in sunlight; clothing is real cotton with weave texture, washed-wear feel and folds with real gravity drape, shoes are real soft leather with wear and clay stuck to the soles. The facial features must be a proper East Asian look — eyes on the narrower/longer side, flat brow bone, clear cheekbones, straight nose bridge; not allowed to be drawn as a Western white person, mixed race or with heavy makeup. The kitten becomes a real short-haired house cat, cream-white with orange-yellow patches, fur individually distinct with fluffy edge light, pink nose tip, pupils constricting with light.
[Lighting] Fully follow the clay video's light direction and shadows: warm sunlight slants from the upper-left-back at about forty-five degrees, shadow edges slightly soft, direction uniform; light shafts contain real floating dust and slight aerial perspective; once colored it carries a clear warm-gold color temperature, shadows falling in clear sky-lit cool blue, forming a natural warm-cool contrast; all AO dark corners of the original clay model are retained as transparent dark areas with ambient color bounce, dark areas have color and detail, not crushed to dead black, highlights not overexposed to white blobs.
[Image quality] Ultra-high-definition 8K realistic cinematography texture, full-frame shallow-to-medium depth of field, natural lens blur and slight chromatic aberration, path-traced global illumination and real contact shadows, high texture resolution with no stretching or tiling, natural motion blur, film-grade color layering and soft highlight roll-off.
Style: realistic live-action cinematic texture, natural-light slice-of-life film, real material and weathering detail. Image quality: 8K ultra-high-definition, full-frame shallow depth of field, path-traced global illumination, film-grade color layering. Lighting: warm sun key light from the upper-left-back paired with clear sky-fill light, long shadows in uniform direction, transparent dark areas with ambient color.
Do not change any camera motion, cut point, aspect ratio, duration, composition or object placement of the original video; do not add or remove any object or character; do not change character bone structure, hairstyle, garment tailoring, body proportions or any action timing; no residual clay areas, locally unpainted pure-white objects, material-lost gray blocks or misaligned textures; no cartoon, anime, illustration, plastic toy or fake CG texture; no Western white faces, mixed-race faces, heavy-makeup faces, face swaps or age jumps; no wireframe grid lines, coordinate axes, UI or any UI elements; no text, Chinese characters, English letters, numbers, labels, subtitles, watermarks, logos or platform corner bugs; no flickering, material jitter, texture swimming, clipping, limb distortion, extra fingers, noise or blown-out loss of structure.

File / webpage to video

Upload docx / pptx / xlsx / pdf or paste a webpage link; the model reads the content automatically. The prompt only needs to state clearly who it is for, what style, and what format. For fine control, write the shot structure, voiceover pace and number-animation effects into the prompt.

NoteUploaded files and webpage links must be publicly accessible resources on the internet; the model cannot read content that requires login or is on an intranet.

Reference materialPromptOutput video
Suyan · Light Essence · Creative Proposal.pptx

Based on my proposal document, generate a Suyan Light Essence brand commercial — a story about "three minutes in the morning." The heroine goes from waking up to skincare to heading out, bare-faced throughout, using natural-light texture to prove: good skin doesn't need covering, only kind treatment. No voiceover, no product-efficacy claims throughout; the "bare-faced confidence" brand attitude is conveyed purely through image mood. Keep the protagonist consistent. The final shot ends on the product image.

H1-2026 Cross-border E-commerce Monthly GMV.xlsx
Total duration 22 seconds, 4 segments concatenated in order. The whole film is on a pure-white background, with soft color blocks and simple line charts, Apple Keynote-style minimalist business style, sans-serif dark-gray font. BGM runs throughout, clearly audible volume, steady rhythm and simple melody; sound-effect layer: color-block slide-in whoosh, number-jump keyboard clicks, key-data pop crisp "ding." The voiceover is in English with a bright, confident tone, about 2.5 words per second, with slight emphasis on key numbers.
Segment 1 (0-5s): Centered wide-shot composition on pure white. Gold light particles converge and condense into the title "H1 2026 · Wan"; after the title disperses, two large numbers slide into the center — left $1,580 (Jan GMV), right $3,460 (Jun GMV), with "Monthly GMV Growth" floating below, and a coral label +45.7% popping up at the bottom. Voiceover: "In the first half of 2026, monthly GMV grew from 15.8 to 34.6 million dollars."
Segment 2 (5-11s): Line chart, X-axis Jan-Jun. Three trend lines drawn simultaneously — coral Southeast Asia, cyan-blue North America, amber Europe, with precise values popping up at each monthly node. Voiceover: "All three markets showed consistent month-over-month growth, with no single dip."
Segment 3 (11-17s): Bar chart on the left + data column on the right. Cyan-blue rounded bars rise in sequence corresponding to North America's monthly figures; five month-over-month growth rates slide out on the right (Feb +14.3% → Jun +46.5%), and finally +200% pops in the center. Voiceover: "North America was our growth engine — up two hundred percent in just six months."
Segment 4 (17-22s): Centered-upper donut chart, coral / cyan-blue / amber three arcs corresponding to the three markets' H1 share, slowly rotating. The central number jumps rapidly from 0 to $138.1M, with the subtitle "Three Markets · All Growing · Zero Downturn" fading in below. Voiceover wraps up steadily: "138 million in total — all markets growing, zero downturns."

Baidu Baike · Quantum Mechanics

Turn the quantum-mechanics content from this webpage into an accessible, fun popular-science short video.

Baidu Baike · Wang Wei

Based on this encyclopedia entry on Wang Wei, I want to make a "Who is Wang Wei" animation video for first-grade children.

Line-guided camera movement

On the reference image, mark the camera-movement path and waypoint order with red lines. The prompt then restates the flight trajectory, pass-through relationships, and speed cadence waypoint by waypoint. The model generates a continuous, single-take camera movement that closely follows the red lines.

Reference materialPromptOutput video

Image 1

Scene 2 — Trojan Horse entering the city
[Hard requirements for the base image]
The final footage completely removes all red lines, numeric markers, arrows, and auxiliary marks; no text or watermarks appear in the frame. Surreal photorealistic 8K classical-oil-painting texture, thick brushstrokes and pigment grain clearly discernible, UE5 cinematic-grade lighting render, warm-toned side-backlight of a Greek afternoon, golden-brown dust churning and drifting through the light shafts. The giant white wooden horse, the wooden trailer and rollers, the taut thick hemp ropes, the Trojan stone city wall and round watchtower on the left, and the distant mountain-city complex all retain their original form, proportions, and positions unchanged. Human proportions are natural and undistorted; only the original rope-pulling crowd and the defenders on the wall remain — no extra figures are added. Robes stay in the original ochre-red, indigo, earth-yellow, and off-white.

[Core camera-movement rules]
First-person FPV drone-perspective, single take, one uninterrupted continuous motion throughout, flying along the complete path of red line 1→2→3→4→5 in Image 1, skipping no segment, simplifying no trajectory, changing no order. The opening is not a straight flat flight — it first rises up-left to clear the tow-rope zone, then dives down in a large wave arc. Waypoint 2 must genuinely pass through the space beneath the horse's belly; it cannot skirt around the horse body or sweep past its side. The overall speed-layered cadence is fast → slow → fast.

[Segmented by waypoint]
Segment 1 (waypoint 1 → waypoint 2): high-speed cut-in thrust, a dive that first rises then drops.
The camera takes off at low altitude from the right edge of the frame, beside the spear tips of the cheering spear-bearers, rises sharply up-left at extreme speed, sweeping over rows of taut tow ropes, climbing above the distant mountain-city complex to skim past it, then sharply yaws downward at the position just right of the horse's chest, pressing the nose down and diving along the top of the crowd's heads and raised spears. Arms and banners rush past on both sides, kicked-up dust is torn into long streamers by the airflow, and the nose dives straight for the shadow beneath the horse's belly.
Segment 2 (waypoint 2, passing through the belly): ducking beneath the horse, the speed suddenly slows.
The camera drops to an extreme low, threading in between the horse's four massive pillar-like legs, passing through the shadowed space directly beneath the belly — overhead are the wood-grain planks and chisel marks of the belly underside, on both sides are thick ropes, trailer planks, and iron-banded rollers, below is the dusty ground trampled by countless feet. Light dims abruptly, leaving only golden light bars leaking through the gaps; the sense of scale presses to its maximum.
Segment 3 (waypoint 2 → waypoint 3): slow motion, tilting up and turning sharply.
The camera emerges from the other side of the belly and, near the rear of the horse close to the left city wall, pulls out a large U-shaped rising arc in slow motion, sweeping past the carved texture of the tail and hind legs. The airframe slowly tilts up, and the massive bulk of the stone city wall and round watchtower in the background gradually fills the frame as the camera rises.
Segment 4 (waypoint 3 → waypoint 4): maintaining slow motion, climbing and orbiting along the spine.
The camera slowly climbs along the horse's spine line from back to front, sweeping past the wood seams of the back, the curly carved mane, and the undulations of the neck muscles one by one. Circling birds in the sky scatter in front of the lens; the defenders densely packed on the wall and the distant mountain city drift slowly backward below.
Segment 5 (waypoint 4 → waypoint 5): sudden acceleration, a head-circling whip to finish.
After passing the top of the mane, the camera suddenly accelerates, whipping right into a full closed loop around the horse's head — first darting into the open sky to the horse's right-front, then sharply reversing and pressing back, and finally speeding straight toward the horse's head. The motion stops dead in the final instant.
[Final hold frame]
The camera holds on a close-up position directly in front of the horse's head: the horse's eye, nostril, and chiseled mane dominate the frame, rim light traces a golden outline around the head, and the background is the defocused blue sky, drifting clouds, and distant mountain city.
[Lighting and quality notes]
No cuts throughout, no shot splits, no transition breaks — retain the native FPV wide-angle slight-distortion texture of a single continuous shot. The afternoon golden light always comes from the upper-right of the frame; while passing through the belly, it undergoes one natural exposure transition from bright to dark and back to bright. Dust-particle and airborne-microdust density intensifies with speed, and the oil-painting brushstroke and pigment grain stay consistent across the entire piece — no photo-real plastic sheen appears.

Image 1

Group 1 copy.2555png
[Hard requirements for the base image]
The final footage completely removes all red lines, numeric markers, arrows, and auxiliary marks; no text or watermarks appear in the frame. Surreal photorealistic 8K Eastern xianxia landscape texture, UE5 cinematic-grade lighting render, warm-toned side-backlight diffused by golden dawn, a churning sea of clouds, Tyndall light shafts piercing the rainforest and waterfall mist, golden micro-dust and water-vapor particles suspended in the air. Blue-tile flying eaves, vermilion lacquer columns, white-stone railing panels, suspended floating isles, and the patterned circular altar array all retain their original form and positions unchanged. Figure proportions are natural and undistorted; only the original walking cultivator silhouette remains — no extra figures are added.

[Core camera-movement rules]
First-person FPV drone-perspective, single take, one uninterrupted continuous motion throughout, flying along the complete path of red line 1→2→3→4→5→6, skipping no segment, simplifying no trajectory, changing no order. Waypoint 2 must genuinely pass through the interior of the hall and back out; it cannot detour around or merely sweep past the outer eaves. The overall speed-layered cadence is fast → slow → fast.

[Segmented by waypoint]
Segment 1 (waypoint 1 → waypoint 2): high-speed cut-in thrust.
The camera takes off from the cliff peak at the upper-left of the frame, diving rapidly along the outer eaves of the stacked blue-tile pavilions on the left, threading through the gaps between flying eaves and ancient pine branches. The airframe swings left and right quickly to dodge purlins and pine boughs; clouds and mist are torn open by the high-speed airflow as it dives straight for the door opening beneath the eaves of the second hall below.
Segment 2 (waypoint 2, passing through the hall): hugging the ground to duck inside, the speed begins to slow.
The camera ducks directly into the open door on the front of the second hall, passing through the hall's interior — vermilion columns, dark caisson ceiling, hanging drapes, and the reflective blue-brick floor sweep past the lens on both sides in turn. Light dims abruptly inside; only golden light bars slicing in through side windows cut across the frame. It then exits through the latticed window-door on the other side of the hall, returning to bright daylight.
Segment 3 (waypoint 2 → waypoint 3): slow-motion transition.
After leaving the hall the speed slows further; the camera drops down to the lower edge of the covered bridge, gliding right along the bridge's bottom beam in a slow tracking shot. Above in the frame, the walking cultivator and fluttering robe-sleeves on the bridge are visible; the wooden texture of the bridge body sweeps past one by one, then the camera sweeps past the foreground cliff edge at the bottom-center of the frame.
Segment 4 (waypoint 3 → waypoint 4): low large-arc sweep, maintaining slow motion.
The camera pulls out a large rightward arc along the cliff, skimming low over vegetation, sweeping past the turquoise deep-pool water in the lower-right. Ripples are pressed into the surface by the airflow and the cliff wall reflects in the water; mist splashes onto the lens, then it grazes past the eave-corner and railing panels of the waterside hall by the pool.
Segment 5 (waypoint 4 → waypoint 5): from slow to fast, lateral wall-hugging thrust.
The camera drops low and accelerates rightward along the pool surface, whipping an S-turn to throw the body around, rushing toward the hall at the base of the great waterfall cliff on the right. It passes through the dense mist kicked up by the crashing falls and the rainbow halo, hugging the hall's flying eaves and the cliff rock at high speed, the speed layering up step by step.
Segment 6 (waypoint 5 → waypoint 6): sudden acceleration, vertical climb through the gate to finish.
At the hall at the cliff base, the camera suddenly tilts up and pulls out, climbing straight up along the right-side waterfall cliff and the tiers of suspended pavilions. Rock and pavilions on both sides drop away rapidly; after breaking out of the mist it continues full-tilt up-right toward the backlit golden mountain-gate archway. Golden light shafts gush from the gate opening and hit the lens head-on, the light ratio blowing out instantly.
[Final hold frame]
The camera holds on a low-angle position directly in front of the mountain-gate archway: the frame is filled with the intense golden light pouring from the gate opening, the bracket eaves of the archway and the silhouettes of the mountains on both sides are fully composed within the frame, and the sea of clouds spreads below toward an infinite distance.
[Lighting and quality notes]
No cuts throughout, no shot splits, no transition breaks — retain the native FPV wide-angle slight-distortion texture of a single continuous shot. The dawn golden light always comes from the upper-right toward the mountain gate; during flight the light transitions gradually from backlit to front-lit direct illumination, and while passing through the hall it undergoes one natural exposure transition from bright to dark and back to bright. Particle density of mist, cloud-vapor, and gold-dust intensifies with speed.

Temporal video transfer

Treat the reference video's action, camera movement, and effects as a temporal signal and transfer it to a new character or scene — keeping the rhythm intact while freely swapping subject and environment.

Action reference

Reference materialPromptOutput video
Character reference image

Reference video: replace the skateboarding man in Video 1 with the woman and outfit in Image 1. Note: replace only the character; keep the original skateboarding motion, trajectory, and skatepark background in the video completely unchanged.

Camera-movement reference

Reference materialPromptOutput video

Reference the camera-movement style of the reference video; change the scene to a cup of coffee placed on a rooftop café table.

Effect reference

Reference materialPromptOutput video
Character reference image

Apply the energy effect from Reference Video 1 to the character in Image 1; generate the same effect and action.

Audio-driven

Audio can drive either the beat (dance-to-beat) or the timbre (dialogue transfer), keeping motion, lip sync, mood, and audio precisely in sync.

Reference materialPromptOutput video

Audio 1

Vertical 9:16, full-body shot, fixed and stable camera. A young woman with long black hair, wearing a white floral-embroidered long-sleeve top, a white pleated mini skirt with a black belt, and white platform sneakers, dances in an idol-style routine in a modern minimalist kitchen, following the beat of the fixed audio. The background is light-gray cabinetry, a white island counter, and a dark night-view window; behind the figure is a square fill light forming a rim light, with indoor cool-white lighting.

Choreograph the dance according to the rhythm, speed, and emotional swells of the fixed audio: heavy drum downbeats correspond to crisp, large-range movements and hold poses; gentle passages correspond to soft body waves and hand gestures; the chorus-climax passages are denser and more explosive. Overall it is an idol-style dance, with movements coherent and fluid, full of rhythm and expressiveness, perfectly in sync with the audio beat.

Motion is smooth, limb proportions natural.

Image 1

Male character reference image

Video 1

Audio 1

Seamlessly replace the woman in Video 1 with the male character from Image 1. The character sits by the window, holding a phone to the ear and speaking.

[Timbre and dialogue] Extract the timbre features from Audio 1, and have the character speak the following dialogue: "Really? That sounds great! When are you coming back? I miss you so much."

[Lip sync and expression] The generated speech must precisely drive the character's lip movements to open and close naturally. Speaking is accompanied by natural blinking, slight head nods, and a sense of breathing; the facial micro-expressions must perfectly match the caring, nostalgic emotion in the dialogue.

[Lighting and scene] The figure's face naturally receives mixed lighting from warm indoor light and cool neon light outside the window. Raindrops outside the window slide down continuously and slowly; the coffee on the table gives off a faint wisp of steam; the character and scene edges blend naturally with no cutout artifacts.

[Camera and quality] A 15-second long take, the camera pushing forward extremely slowly, the frame rock-steady, cinematic lighting, 8K resolution, photographic realism.

Video editing

Feed in the original video and use natural-language instructions for precise editing — only the specified content changes; the rest of the frame stays unchanged. A reference phrasing is Edit the video: change xxx in video 1 to yyy, keep the rest of the frame unchanged (the leading "edit the video" can be omitted). Use the verbs add / delete / replace / change to / reshape; modifying one item at a time is the most stable.

  • Element editing: Describe the elements to add, remove, or modify. When reference images are involved, refer to them by name as image 1 / image 2.
    • Add: The man puts on a red hat.
    • Modify: Replace the cat with a dog.
    • Delete: Remove the cat on the grass.
    • Reference image: Have the man in the video wear the top from image 1 and the pants from image 2.
  • Global parameter editing: Describe changes to the overall style, lighting, color tone, or weather. When the rest should be preserved, add the phrase keep everything else unchanged.
    • Weather: Change the video to overcast weather.
    • Color tone: Change to a warm color tone.
    • Style: Change the video to a 2D cartoon style.
  • Temporal editing: Describe how actions / dialogue change over time. You can pair this with 4-6s timestamps to pinpoint a segment.
    • Action: Have the man in the video pick up a water glass and take a sip of water.
    • Dialogue (preserving voice timbre): Keep the character's voice timbre and tone; change the woman's spoken line to: 'Hello World'
TypeVideo before editingPromptOutput video
Add element

Video 1

Edit the content of video 1. Add a reef on the beach, half-buried in the wet sand; when the tide recedes, let the seawater gently wash over the base of the reef.

Modify element

Video 1

Edit the video: replace the skateboarding man in video 1 with a short-haired woman wearing a sports tank top and loose cargo pants. Keep the baseball cap, knee pads, elbow pads, and other protective gear from the original video. The skateboarding motion, movement trajectory, and skatepark background stay exactly the same.

Delete element

Video 1

Edit the video: remove the woman's sunglasses in video 1, revealing natural eye makeup and gaze; the woman's movement and the rest of the frame stay unchanged.

Lighting edit

Video 1

Edit the video: brighten the overall lighting of video 1, applying soft three-point lighting with a TV-drama feel. Eliminate the hard shadow on the right side of the man's face so the facial features are evenly lit, with natural, translucent skin tone. Synchronously boost the light on the white tiled wall and window in the background, keep the overall image tone consistent, and leave the man's movement and frame content unchanged.

Action edit

Video 1

Edit the video: swap the positions of the suited man and the ostrich in video 1. Move the ostrich to a forward position on the left side of the frame, and move the man to the right side right behind the ostrich. Keep both figures' walking motion and stride rhythm consistent; the street background stays unchanged.

Style edit

Video 1

Edit the video: change the video style to claymation style.

Dialogue edit

Video 1

Edit the video: change the young man's dialogue in the video to: "The deal is done. Now... we disappear."

Reference edit

Video 1

Image 1

Reference image 1: 10.png

Image 2

Reference image 2: light two-tone baseball cap

Image 3

Reference image 3: denim shirt

Edit the video: have the woman in video 1 put on the hat from image 1, naturally fitting her head shape. Have the man in video 1 put on the hat from image 2, naturally fitting his head shape; replace the man's ochre-brown shirt with the blue washed loose denim shirt from image 3, collar unbuttoned and sleeves rolled up to the forearms. Both figures' movements, clothing, and the rest of the frame stay unchanged.

Plot reshape

Video 1

Reshape the plot of video 1: have the man directly pick up the guitar and begin playing a piece of music.

Video extension

Continue the footage from the original video; the style, subject, and scene carry over from the original. The prompt should clearly state the extension intent, direction, and new content:

  • Extension intent: The prompt should contain keywords such as extend / continue / continue writing.
  • Extension direction: Extend backward (starting from the last frame), Extend forward (ending at the first frame), Extend forward and backward (using the original video as the middle segment).
  • Extension content: Clearly state the new actions, dialogue, and frame changes, and restate the character traits, clothing, and scene to avoid continuity glitches in the continued segment; you can use "timestamp + shot" to carry over the mood and camera movement from the original.

NoteInput video duration + output video duration ≤ 30 seconds. It is recommended to set duration to -1, or to the sum of the input duration and the expected extension duration.

TypeVideo before editingPromptOutput video
Extend backward

Extend video 1 by 15s. Character A is a man with slicked-back hair, wearing a dark-gray Chinese-style long gown, with a brown prayer-bead bracelet on his left wrist; character B is a man wearing a dark Chinese-style long gown. Character B slowly sets down the teacup, his gaze dropping slightly, and taps the tabletop twice with his fingertips; unhurriedly he says, "The price is easy to discuss, but when will the goods arrive at port? You need to give me an exact date." Character A, his smile undiminished, picks up the lidded teacup and skims the tea foam. Character A takes a sip of tea, sets the teacup back on the saucer, leans forward and lowers his voice; the prayer beads sway gently with the gesture as he says in a firm tone, "By the end of the month at the latest. I've already sorted everything out at the dock — a full shipment will be delivered neatly to the door of your warehouse." With that, he presses his right hand palm-down on the tabletop, his gaze locked on character B. Character B mulls for a moment, his eyes sweeping to the "Wan An Dang" plaque outside the window; the bamboo curtain lifts gently in the street breeze. He turns back, a smile creeping onto his lips, picks up the lidded teacup and offers a slight toast toward character A: "Good — delivery at the end of the month, I'll wait for your good news." The two smile at each other. Keep the steady, grounded tone of a Republican-era business-war film; the dialogue rhythm is unhurried, with low tea-house ambient sound underneath.

Extend forward

Extend video 1 forward by 15s. Character A is a dark-brown short-haired man wearing a black tailcoat, white shirt, black bow tie, and beige waistcoat; character B is a brown-haired woman in an updo, wearing a beige puff-sleeve vintage long dress with black lace trim and dark teardrop earrings. The crystal chandelier in the ballroom casts light over the marble floor; guests converse in twos and threes, holding glasses and speaking in low tones, as a melodious waltz string prelude slowly begins and the crowd's chatter gradually quiets. Character B walks slowly from one side of the ballroom, her hem brushing the marble floor, her gaze sweeping the room. Character A steps out from the crowd, threading through the guests straight to character B, gives a slight bow, extends his right hand, and gazes at her tenderly. Character B lowers her eyes with a faint smile, rests her right hand lightly on character A's palm; the two share a look and a smile. Character A leads character B to the center of the dance floor; the surrounding guests naturally step back to clear the space, all eyes on the two of them. Character A holds character B's right hand raised with his left hand, his right hand lightly resting on the small of her back; the two assume the waltz posture, step out on the first beat with the strings, and begin to spin. An elegant 19th-century palace-ball atmosphere, warm golden tone.

Extend forward and backward

Using the video as the middle segment, extend forward by two seconds: the camera follows smoothly as the girl walks slowly toward the lens, stops, gently tilts her head up, takes a deep breath in the cold air, lips parting slightly; using the video as the middle segment, extend backward by three seconds: after rubbing her hands, the girl looks straight into the lens and breaks into a warm, radiant smile. She slowly raises one gloved hand to catch a gently falling snowflake; the camera slowly pulls back into a medium-wide shot, revealing the sunlit snowy forest path, golden-hour backlight, lens flare, shallow depth of field, warm and healing atmosphere.

First frame / First and last frame

Use a first frame (or first frame + last frame) to lock the start and end of the picture; the model is responsible for completing the motion in between. The image already establishes the subject, scene, and style, so the prompt should focus on describing the motion, plot, dialogue, and camera movement.

NoteFirst-and-last-frame mode and all-reference mode are

mutually exclusive; once enabled, reference material input is no longer supported.

First frame to video

Input first framePromptOutput video
Ink bamboo grove first frame
(0:00-0:03) Camera movement: medium shot.
Picture: Continue the ink-wash bamboo grove scene from image 1. A swordsman wearing a bamboo hat and a dark robe stands as a silhouette in the center, surrounded by thick mist and swaying bamboo shadows.
Action: The swordsman rests his left hand lightly on the hilt; bamboo leaves drift down in the wind.
Mood: Silent, oppressive; the granular texture of the ink wash is clearly visible.
(0:03-0:08) Plot introduction, camera movement: extreme close-up.
Plot: The swordsman senses killing intent.
Picture: The camera quickly pushes in to an extreme close-up beneath the swordsman's bamboo hat.
Details: Beneath the hat brim, a pair of sharp, resolute eyes is revealed, the gaze like a blade. Ink-wash sweat beads slide down from his forehead.
Action: The swordsman slowly grips the hilt with his right hand, knuckles turning white. The falling bamboo leaves accelerate.
(0:08-0:13) Camera movement: rapid orbit.
Picture: The camera orbits rapidly around the swordsman as the center.
Plot: Two assassins, also rendered as ink-wash silhouettes, burst from the depths of the bamboo grove, wielding twin hooks and a short blade.
Details: The assassins' movements are likewise formed by ink-brush strokes, flowing like ink traces. The camera sweeps swiftly past the assassins' figures, emphasizing a sense of speed.
(0:13-0:18) Camera movement: handheld feel, high-speed capture.
Combat: The swordsman suddenly draws his blade; the flash of the blade is a pure-white ink-wash fissure. The assassins attack with twin hooks.
Details: First strike — the swordsman turns and slashes; blade meets the assassin's twin hooks, bursting into countless black-and-white sparks and ink splatters with a calligraphic feel. Second strike — the swordsman feints, springs off a bamboo pole, and strikes down from above; the assassin tumbles away, the bamboo pole snaps and scatters into flying ink marks. Third strike — the swordsman's blade tip points straight at another assassin's throat; the assassin parries with a short blade, and the camera captures the fine ink-trace fracturing detail as blade and short blade grind together.
(0:18-0:20) Camera movement: slow motion to freeze-frame.
Picture: The swordsman's blade stops beside the assassin's neck; the assassin freezes.
Details: The swordsman's bamboo hat is slightly tilted from the fight. The bamboo grove returns to calm; more bamboo leaves flutter.
Mood: The camera slowly pulls back to a medium shot, freeze-framing on the swordsman sheathing his blade and the assassin collapsed on the ground.
Text (optional): Ink-wash-style text appears at the bottom of the frame: "The jianghu is but a single stroke of erasure."
Style notes: Maintain the high-contrast black-and-white ink-painting style of image 1 throughout. The action must carry a brush-stroke feel, the speed must be fast, and the details rich (ink drops, sparks, ink-trace fissures).

First and last frame to video

First frame & last framePromptOutput video
Woman in garden first frameWoman in garden last frame
Shot 1 (00:00-00:08) Beautiful past:
35mm film stock, a vintage sunlit garden. An elegant woman wearing a straw hat, a black necklace, and a white dress stands before a wall of blooming roses. The wind stirs her hair; she lightly touches the hat brim, her gaze tender with a trace of contemplation. The camera slowly pushes in, natural light flares, soft warm tones, shallow depth of field, exquisitely delicate motion, 4K high quality.
Shot 2 (00:08-00:15) Turn of mood:
Cinematic close-up. The woman slowly removes the straw hat from her head and lowers her gaze, her expression carrying a faint wistfulness and longing. A soft breeze passes; the light and shadow shift subtly, conveying a story-like narrative atmosphere. 35mm film texture, deep and elegant, cinematic lighting.
Shot 3 (00:15-00:23) Return to reality — inside the car:
View from the back seat of a moving car, the same woman in profile, her hand resting on the door, gazing out the rain-streaked window. Blurred greenery and dappled light outside rush backward; the light inside the car is dim yet warm. Her profile is refined, highly cinematic in emotional feel, authentic film grain, slow cinematography, elegant and melancholic.
Shot 4 (00:23-00:30) The final gaze:
Extreme facial close-up. Beside the car window, the woman slowly closes her eyes, revealing a calm, released expression. Raindrops glide slowly down the window glass; the focus gradually shifts onto the raindrops on the glass, the background figure blurs, and the picture slowly fades to black. Wong Kar-wai-style cinematic feel, deeply emotional and fluid.
Mineral-color Xizi first frameMineral-color Xizi last frame
[Single continuous take, slow and aesthetic camera movement, no cuts, 15-second long shot, collision of mineral-pigment rock-color painting texture with 3D rendering, teal-blue × gilded-gold color clash]
Global setting: Chinese-style rock-color painting texture, mineral-pigment granular grain, gold-leaf texture, a strong color clash between teal-blue and gilded gold. The scene is West Lake in morning mist at dawn: the lake surface is a matte teal-blue wash, golden morning light pierces the clouds to form Tyndall light shafts, and the stone pagodas of Three Pools Mirroring the Moon and Leifeng Pagoda loom faintly in the mist, with lotus leaves and flowers nearby. The main subject is the West Lake water goddess "Xizi": rendered in rock-color texture, oval face, willow-leaf eyebrows, phoenix eyes, a vermillion floral mark at the center of her brow, skin like lustrous white jade with a soft glow, and a pale-gold halo suspended behind her head; her hair is coiled high and pinned with a gilded lotus-flower hairpin, and her skirt and sash are formed from flowing lake water, a teal-blue gradient with shimmering ripples on the surface.
Voiceover and sound effects: The background features a gentle guzheng and bamboo flute; ambient sound holds subtle ripples and dawn birdsong. From the 8th second, an ethereal, tender female voice slowly recites "If one were to compare West Lake to Xizi —"; during the glance-back segment it continues with the second half "Lightly adorned or richly made up, always fitting." The pacing is slow, with a faint reverberation and lingering resonance.
[0-3 seconds: A single drop of water]
Extreme macro. At the edge of a lotus leaf at dawn, a water drop gathers full, reflecting the golden morning light, and slowly falls. The camera follows the drop downward (Tilt down); the drop strikes the lake, setting off a ring of gold-blue ripples that expand outward in layers.
[3-7 seconds: Water gathers into form]
At the center of the ripples, the lake water slowly rises and condenses into the figure of the goddess "Xizi": the water becomes her teal-blue skirt and translucent sash, and tiny streams flow around her. She rises slowly from the heart of the lake, her eyes gently closing then slowly opening, the gilded lotus hairpin glowing faintly. The camera slowly orbits (Orbit) her half-lowered eyes and the skirt woven from flowing water.
[7-11 seconds: A lotus with every step]
She walks barefoot and slowly across the lake surface; with each step, a ring of golden ripples spreads beneath her feet, and lotus flowers bloom in succession within the ripples. The camera pulls back to a medium-wide shot: the thin mist parts, the three stone pagodas of Three Pools Mirroring the Moon light up warmly in the distance, Leifeng Pagoda emerges in the morning light, and weeping willows brush the water along the shore. At the 8th second, the female voiceover begins "If one were to compare West Lake to Xizi —"; her sash trails a long streak of gold-blue water-light behind her.
[11-15 seconds: The glance back — always fitting]
She stops and glances back toward the camera; the camera slowly pushes in to a half-body close-up: morning light falls on her face, the lake-water sash drifts gently at her side, and a few glowing water drops fall from the tips of her hair, dissolving into fine points of light in midair. The voiceover continues with the second half "Lightly adorned or richly made up, always fitting." She smiles faintly; behind her are Leifeng Pagoda and the misty Three Pools Mirroring the Moon in the morning light. In the final second, the image fully freezes on her glance-back close-up, locking into a perfect cinematic-poster composition.

Industry best practices

Short drama

Reference materialPromptOutput video

Image 1

Image 1: protagonist Shen Yan

Image 2

Image 2: antagonist Lu Feng

Image 3

Image 3: abandoned factory location
Protagonist: Image 1
Antagonist: Image 2
Scene: Image 3 — inside an abandoned factory, afternoon, side-back light streaming through high windows, dust floating in the air, puddles and broken glass on the ground, two rusty iron pillars in the midground, wooden crates and empty glass bottles stacked to the rear right, a rusted iron barrel lying on its side to the left, a few iron chains hanging from the ceiling beams.
Style: realistic combat.
Character relationship: The protagonist is Shen Yan, a Chinese woman, twenty-eight years old, lean and wiry, hair in a low ponytail, wearing a grey-blue quick-dry long-sleeve top and dark-grey cargo pants, both hands wrapped in white boxing hand wraps. She excels at close-quarters clinch fighting, throws and takedowns, and joint locks, making up for her strength disadvantage with angles, leverage, and weight transfer. The antagonist is Lu Feng, a Chinese man, forty years old, tall and heavy-set, wearing a black tank top, sandy-brown cargo pants, and high leather boots, wielding a sixty-centimeter rusty steel pipe in his right hand. He holds an absolute advantage in strength and weight, suppressing with the pipe as a long weapon and brute force. Throughout the film the difference in build must be recognizable at a glance, and the protagonist never wins by matching force with force.
General action-design principles: The whole film is one unbroken continuous combat chain. Each shot's ending posture is the next shot's starting posture. Character positioning, weight direction, and force direction must carry over — no teleporting, no inexplicable repositioning, no movements that violate center of gravity or joint range of motion. All throws and joint techniques must follow the real biomechanical sequence of jiu-jitsu and wrestling: first break the opponent's balance, then apply leverage along the direction of their momentum. Never depict superhuman feats such as lifting or throwing a much heavier opponent off the ground with one hand.
Shot 1
Duration: 1.4s
Shot size: medium close-up
Camera angle: eye level
Camera movement: slow lateral dolly + breathing micro-sway
Focal length and depth of field: medium depth of field
Light: side-back light, high contrast, warm orange + cool cyan
Frame description: The two stand two meters apart, facing off across a slanted shaft of light. Shen Yan lowers her center of gravity, left foot forward in a bladed fighting stance, both hands guarding between her cheeks and jaw, knuckle wraps taut. Lu Feng stands with feet apart, right hand holding the pipe hanging down so the barrel points diagonally at the ground, shoulders and back rising and falling with his breath. Collision point: none; shockwave form: none; medium layer: dust in the light shaft rolls slowly, the white breath from both is visible in the backlight; destruction layer: the puddle surface ripples faintly from the slight grinding of foot soles; lens-feedback layer: extremely slow lateral dolly with a slight breathing float, focus held sharp throughout.
BGM: low sustained string drone, layered with a slow heartbeat sub-bass.
Voice/sound effects: the two's interleaving breathing (center), distant water drips (surround), short scraping of the pipe end against the floor (right channel).
Chinese AI prompt: realistic film, medium close-up, an Asian female fighter and a tall Asian man with a steel pipe face off inside an abandoned factory across a side-back light shaft, floating dust, high-contrast warm orange and cool cyan, HD realistic photography, 8K, cinematic lighting.
Camera-movement instruction: extremely slow lateral dolly from left to right with a slight breathing up-down float, the last frame held on the composition of the two facing each other directly.
Shot 2
Duration: 1.6s
Shot size: medium shot
Camera angle: eye level tilted right, slight Dutch angle
Camera movement: rapid push-in + small whip-pan following the attack arc
Focal length and depth of field: shallow depth of field
Light: side-back light, hard light
Frame description: Lu Feng steps in to press forward, the pipe chopping hard from upper right to lower left. Shen Yan does not block hard — she steps diagonally forward-left, her head tilting aside, the pipe grazes past her right ear and slams into the rusted iron barrel beside her; in the same instant her right hand has already reverse-gripped Lu Feng's right wrist that holds the pipe. Collision point: pipe against barrel wall; shockwave form: the barrel wall dents inward at the impact point and a visible ring of air ripple spreads outward; medium layer: the air ripples distort, dust accumulated inside the barrel puffs out the opening in a column; destruction layer: the barrel jumps sideways half a meter from the impact, a fresh dent left on its body, broken glass on the ground jumps from the shock; lens-feedback layer: the impact frame judders violently sideways once, loses focus for 0.08s, slight purple-edge chromatic aberration at the frame edges.
BGM: percussive accent punches in, strings shift from sustained drone to rapid staccato.
Voice/sound effects: the massive metallic crash of pipe on barrel and its buzzing resonance (center), the muffled poof of dust ejecting (left channel), Shen Yan's short inhale (near field).
Realistic combat, medium shot, a tall man swings a rusty steel pipe downward in a chop, an Asian woman side-steps to evade while reverse-gripping his wrist, the pipe strikes the rusted barrel sending dust and a dent, lens judder and defocus, cinematic lighting, 8K.
Camera-movement instruction: rapid push-in, small whip-pan following the pipe's downward arc, violent lateral judder at the impact instant then rapid stabilization.
Shot 3
Duration: 1.7s
Shot size: close-up
Camera angle: slight low angle
Camera movement: fixed position + short upward tilt following the motion
Focal length and depth of field: shallow depth of field
Light: side-back light, dirty-lens dust
Frame description: Continuing the wrist grip, Shen Yan drives the heel of her left palm upward into the inside of Lu Feng's elbow on the arm holding the pipe, while her right hand twists his wrist outward — a standard weapon-disarm force sequence: first pry the elbow to strip the arm of support, then twist the wrist to force the tiger's mouth open. Lu Feng's fingers are forced open, the pipe slips free and tumbles away toward the upper left of frame, the barrel rolling one and a half turns through the light shaft. Collision point: palm heel against the elbow pit, wrist joint; shockwave form: a small ring of air pulse at the struck elbow; medium layer: dust in the air is churned into turbulence by both arms, sweat beads on Lu Feng's forearm shaken off as dots; destruction layer: Lu Feng's entire right arm is pushed up, his shoulder line skews, his center of gravity forced to lean forward half a step; lens-feedback layer: a short upward tilt following the elbow strike, fine dust settles on the lens glass creating a dirty-lens effect, no defocus.
BGM: rapid dense drum rolls, viola joins with a fast ascending run.
Voice/sound effects: the dull meaty impact of the palm heel on the elbow pit (center), the faint crack of the wrist joint being twisted (near field left), the metallic howl of the pipe tumbling free (panning from center to upper left).
Realistic film, close-up, slight low angle, an Asian woman strikes the opponent's elbow pit upward with her palm heel and twists the wrist to disarm, the rusty steel pipe flies free tumbling, sweat beads scatter, dirty-lens effect, cinematic lighting, 8K.
Fixed position, short upward tilt following the elbow-strike force, dirty-lens maintained, last frame still has the pipe in midair.
Shot 4
Duration: 1.8s
Shot size: medium shot
Camera angle: eye level
Camera movement: close-body orbiting tracking shot (about ninety degrees)
Focal length and depth of field: medium depth of field
Light: side-back light, the light shaft passing between the two
Frame description: No time for the opponent to recover. Following the already-twisted direction, Shen Yan presses and rotates Lu Feng's right arm further outward along his body, while her left foot hooks behind his right foot to trip his support leg and her right knee drives into the outside of his thigh — his center of gravity is pulled out from underneath, Lu Feng's upper body is forced to pitch forward and down, one hand reaching toward the puddle on the ground. The lens orbits close to their bodies by nearly ninety degrees to lay out the spatial relationships of the action clearly. Collision point: knee against outside of thigh, ankle against ankle; shockwave form: no burst point, only clothing creases and air flow from the bodies compressing; medium layer: their rapid turn kicks up ground dust into a ring, the light shaft is cut by a body and flickers once; destruction layer: the puddle is smeared into an arc of water by the shoe sole, broken glass is kicked aside with a fine tinkling, Lu Feng's bracing palm presses a handprint into the ground; lens-feedback layer: the image stays stable during the orbit, with only one extremely brief stutter on the frame where the leg-trip takes load.
BGM: strings sustain a rising line, percussion switches to syncopated rhythm.
Voice/sound effects: the friction of their clothing and fabric pulling taut (near field), the squeak of shoe soles slipping on wet ground (below), Lu Feng's roar (center).
Realistic combat, medium shot, orbiting tracking shot, an Asian woman pins the opponent's arm against its joint and trips his leg to break his balance, the tall man pitches forward bracing with one hand, the puddle is smeared into an arc of water, floating dust orbiting, cinematic lighting, 8K.
Close-body orbiting tracking shot about ninety degrees, one extremely brief stutter on the leg-trip load frame, last frame held on Lu Feng bracing on the ground with Shen Yan above and to his side.
Shot 5
Duration: 1.9s
Shot size: medium close-up
Camera angle: eye level transitioning to slightly high
Camera movement: handheld tracking + violent shake + short pull-back
Focal length and depth of field: shallow depth of field
Light: side-back light
Frame description: The situation reverses. Lu Feng ignores his arm and uses his weight and brute force to heave himself up, reverse-grabbing Shen Yan's left forearm and flinging her whole body to the side and back; her feet are dragged half a step off the ground and her back slams hard into the midground rusty iron pillar. Riding the rebound of the impact, she pushes off the pillar with both legs, core engaged, and rotates half a turn around the arm still being gripped — this is leveraging force, not a jump, using the gripped arm as an axis. Collision point: back against the pillar, soles of both feet against the pillar face; shockwave form: pillar vibration and dust rings spreading up and down from the impact point; medium layer: accumulated dust on the pillar shakes off into a falling dust curtain, fine debris in the air lit by the backlight; destruction layer: the pillar visibly vibrates and sheds rust flakes, water pooled at its base is shaken into ripples, the clothing on Shen Yan's back is scuffed grey; lens-feedback layer: the impact frame judders violently twice, loses focus for 0.12s, then a short pull-back to accommodate her rotation arc, dirty-lens worsens.
BGM: the music abruptly drops out for 0.2s, leaving only a single metallic overtone, then the full weight of all instruments punches back in.
Voice/sound effects: the heavy dull thud of the back hitting the pillar and the pillar's buzz (center), Lu Feng's low growl of exertion (right), Shen Yan's muffled grunt as she is thrown (near field), the fine rustling of rust flakes landing (below).
Realistic film, medium close-up, a tall man grabs an Asian woman's forearm and flings her into a rusty iron pillar, she kicks off the pillar with both feet and rides the rebound to rotate half a turn around the arm, the pillar sheds a large amount of rust and dust, backlight dust curtain, violent lens judder and defocus, dirty lens, cinematic lighting, 8K.
Handheld tracking with violent shake, the pillar-impact frame judders twice and defocuses, then a short pull-back follows the rotation, dirty-lens maintained.
Shot 6
Duration: 2.1s
Shot size: medium shot
Camera angle: slow descent from eye level to slight low angle
Camera movement: drops in sync with the falling body + small orbit
Focal length and depth of field: medium depth of field
Light: side-back light, puddle reflections on the ground
Frame description: The instant the rotation completes, Shen Yan's right arm has already hooked around the side of Lu Feng's neck from behind, her legs crossed and locked tight around his waist from behind, her full weight hanging on his back pulling downward — not a static choke hold, but using her own body weight to continuously drag the opponent's center of gravity backward. Lu Feng stumbles two steps, both hands clawing behind him, and finally falls onto his back, slamming into the puddle and broken glass on the ground, water and glass shards exploding outward simultaneously. The lens drops together with the falling body, the camera settling to near-ground height. Collision point: Lu Feng's back against the ground, back of his head against the puddle; shockwave form: a circular splash crown spreading from the impact center, plus a fan-shaped ejection of broken glass; medium layer: the splash glows as golden particles in the backlight, the air is pressed into a visible ring of ripples; destruction layer: the puddle is blasted open into a water crown one and a half meters across, broken glass ejects outward and traces fine water lines, the ground dust is pressed flat then curls outward again; lens-feedback layer: the landing frame vibrates violently, loses focus for 0.15s, water droplets splash onto the lens glass creating an obvious dirty lens.
BGM: brass sustains a descending top-note, timpani strikes twice in succession.
Voice/sound effects: the heavy dull thud of the body hitting the ground (center), the crash of the splash (full surround), the crisp tinkling of broken glass ejecting (left and right), Lu Feng's heavy ragged struggling breaths as his airway is compressed (near field).
Realistic combat, medium shot, an Asian woman locks the tall man's neck from behind with her arm and clamps his waist with her legs, dragging the opponent backward to the ground with her own weight, the man falls onto his back into the puddle, splash and broken glass burst radially, backlit golden water droplets, the camera sinks with the fall, water-splashed dirty lens, cinematic lighting, 8K.
Drops in sync with the falling body to a near-ground camera position, small orbit keeping the neck-lock relationship readable, violent vibration and defocus on the landing frame, then dirty lens with water droplets.
Shot 7
Duration: 2.3s
Shot size: close-up → medium shot
Camera angle: near-ground low angle
Camera movement: fixed near-ground position + rapid small push-in after the hit
Focal length and depth of field: shallow depth of field
Light: side-back light, ground reflection lighting the jaw
Frame description: Ground grapple. Lu Feng wrenches his arm free from the neck lock with brute force, rolls and reverses, pinning Shen Yan with his weight and raising his right fist; the instant the fist falls, Shen Yan's left hand feels a shard of broken glass in the puddle, and she does not stab with it — instead she uses the glass to reflect the window light, a sharp white light-spot shooting straight into Lu Feng's eyes. He instinctively shuts his eyes and turns his head, the fist swings wide and slams into the puddle beside her ear. In that half-second, Shen Yan raises both legs to trap one of his arms and his neck, hips thrusting upward to complete the armbar lock, Lu Feng's elbow joint pulled to its limit, his whole body forced to roll to the side. Collision point: fist against the puddle, legs against the elbow joint; shockwave form: a columnar splash of water bursts upward from the fist impact; medium layer: the splash and the light-spot simultaneously form a scattering halo in the air, dust lit up by the light-spot; destruction layer: the puddle is punched into a crater and radial water ripples, broken glass is pushed aside by the fist wind, the skin on Lu Feng's elbow is pulled into taut creases; lens-feedback layer: as the light-spot sweeps across the lens a horizontal lens flare and slight glare appear, the fist-impact frame vibrates once and loses focus for 0.1s, then a rapid small push-in to the detail of the locked elbow joint.
BGM: all instruments resolve into a single sustained high-frequency taut string note, percussion lands one heavy hit as the fist falls.
Voice/sound effects: the heavy dull water-sound of the fist plunging into the puddle (center), a high-frequency metallic overtone as the glass reflects light (stylized foley, right), Lu Feng's heavy gasp and muffled roar as his elbow is wrenched taut (near field), Shen Yan's short exhale of exertion (near field).
Realistic film, near-ground low angle, ground grapple, a tall man reverses and throws a punch, an Asian woman uses a shard of broken glass to reflect window light into his eyes, making him swing wide into the puddle, then clamps his arm with her legs to complete an armbar locking his elbow joint, splash bursts, lens flare, shallow depth of field, cinematic lighting, 8K.
Fixed near-ground position, a horizontal lens flare as the light-spot sweeps, vibration and defocus on the fist-impact frame, then a rapid small push-in to the locked-elbow detail.
Shot 8
Duration: 2.2s
Shot size: wide shot
Camera angle: low angle → eye level
Camera movement: rapid pull-back + slow rotation back to position
Focal length and depth of field: deep depth of field
Light: side-back light, puddle highlight reflections on the ground
Frame description: Resolution and loop closure. In agony, Lu Feng pushes himself up with his other hand, lifting Shen Yan's whole body with him, and the two rise together; she releases the leg lock riding that upward force, both feet planting back on the ground, while her right foot precisely side-kicks the middle of the steel pipe lying on the ground — the pipe is kicked into a high-speed spin on the ground and ejected, slamming horizontally into his shin, he loses his balance and pitches sideways into the stack of empty glass bottles to the rear right, the bottles shatter in sheets, shards fan out at high speed, and after falling he slides another meter before stopping. Shen Yan shakes her numb left hand, walks slowly back to the original position from Shot 1, lowers her center of gravity again and raises both hands to guard her face; the last frame has her posture, position, and center-of-gravity height completely identical to the first frame of Shot 1, dust and the last few glass shards settling, completing the loop preparation. Collision point: sole against the pipe, pipe against the shin, body against the bottle stack; shockwave form: the pipe's ejection traces a low air-flow line, the bottle-stack shatter bursts out in a fan; medium layer: the air drives broken glass and dust outward, forming a whole curtain of glittering particles in the backlight; destruction layer: all the glass bottles shatter and cascade off the wooden crates, the ground is scored by a wet streak from the body's slide, the pipe bounces twice more on the ground after ejecting, the hanging chains in the distance sway gently in the air flow; lens-feedback layer: the shatter frame vibrates violently, loses focus for 0.18s, then a rapid pull-back and slow rotation return to the Shot 1 camera position and composition, the dirty lens fading out with the last breath.
BGM: the full band lands one tutti closing accent, then rapidly decays into the viola's long-tail resonance, finally leaving only ambient room tone.
Voice/sound effects: the high-frequency metallic buzz of the pipe ejecting (panning from center to right), glass bottles shattering in sheets (full surround), the dull thud and friction of the body falling and sliding (right), Shen Yan's heavy breathing gradually steadying (near field), distant water drips re-emerging (surround).
Realistic film, wide shot, low angle, an Asian woman side-kicks the steel pipe on the ground to send it spinning into the opponent's shin, the tall man loses balance and crashes into a stack of empty glass bottles, shards fan out, then the woman walks back to her original position and raises her hands again, the last frame's composition matches the first, dust and glass settling, cinematic lighting, 8K.
Camera-movement instruction: violent vibration and defocus on the shatter frame, then a rapid pull-back and slow rotation back to the Shot 1 camera position and composition, stable finish, dirty lens fading out.

Image 1

Image 1: Ye Xingchen

Image 2

Image 2: female inner-sect disciple

Image 3

Image 3: exterior of the cave dwelling

Image 4

Image 4: interior of the cave dwelling
Total video duration: 15 seconds
Aspect ratio: 16:9
Art style: CG film style
Core plot:
Continuing from the last frame of the previous segment, Ye Xingchen returns to his own cave dwelling and closes his eyes to regulate his breath; a night passes. The next morning, the sect bell tolls distantly, mountain mist still lingering. Ye Xingchen opens his eyes inside the cave dwelling; the Foundation Establishment aura has been suppressed back into his body, and his appearance still resembles a late Qi Refining disciple. A knock comes from outside the cave dwelling, and a female inner-sect disciple arrives to pass on a message: Senior Sister Su has asked him to gather at the mountain gate at the Chen hour tomorrow for the Wanbao Trading Meet, held once every three years at the Holy Land's center. Ye Xingchen's surface stays calm, his heart stirring faintly.
Dialogue:
Female inner-sect disciple: "Junior Brother Ye, Senior Sister Su asks you to gather at the mountain gate at the Chen hour tomorrow."
Ye Xingchen: "For what matter?"
Female inner-sect disciple: "The Wanbao Trading Meet, held once every three years at the Holy Land's center, is about to open."
Female inner-sect disciple: "The sect is bringing a few disciples along to see the world this time, and Senior Sister Su named you."
Image references:
Image 1 is Ye Xingchen.
Image 2 is the female inner-sect disciple.
Image 3 is the exterior of the cave dwelling.
Image 4 is the interior of the cave dwelling.
Fixed spatial relationships:
Opens in the cave interior, Image 4, with Ye Xingchen sitting cross-legged at the center of the stone chamber or before the stone bed. Mid-segment moves to the cave doorway, with Ye Xingchen standing inside to the left and the female inner-sect disciple standing outside to the right or at the foot of the steps, the exterior being the cave dwelling exterior, Image 3. The two speak face to face, Ye Xingchen on the left and the female inner-sect disciple on the right; their clothing, hairstyle, face shape, and bearing must be clearly distinct. The last frame holds on Ye Xingchen standing at the cave doorway, his eyes shifting faintly after hearing the news.
0-2s:
Continuing from the last frame of the previous segment. Ye Xingchen sits cross-legged in the cave interior; the dwelling is simply furnished, with the stone bed, stone table, clear-water pool, or cultivation area quietly visible. Faint morning light filters in through the door gap or stone window, falling on his sleeve and the side of his face. In the distance the sect bell tolls, mountain mist still lingering.
2-4s:
Ye Xingchen slowly opens his eyes. The lens gives a close-up of his face, with realistic skin texture, eyelashes, and clear reflections in his eyes. His breathing is steady, shoulders and neck relaxed, the spiritual fluctuation from Foundation Establishment fully withdrawn, his appearance still resembling a late Qi Refining disciple.
4-6s:
A soft knock comes from outside the cave dwelling. Ye Xingchen looks up toward the stone door, rises, and adjusts his sleeves and collar. The lens gives close-ups of his hands, cuffs, and the storage pouch at his waist, showing him tucking away any abnormal aura, his movements restrained and quiet.
6-8s:
The lens cuts to the cave doorway. Ye Xingchen opens the stone door, and morning mountain mist drifts slowly past outside. The female inner-sect disciple stands outside to the right, posture respectful. She cups her hands and says: "Junior Brother Ye, Senior Sister Su asks you to gather at the mountain gate at the Chen hour tomorrow." Her face is soft and refined, clearly different from Ye Xingchen's cold, lean face.
8-10s:
The lens cuts to a half-body close-up of Ye Xingchen. He stands inside to the left, his expression calm, not showing much emotion, only asking: "For what matter?" The morning breeze gently moves his sleeves, the cave's stone door and mountain mist staying sharp in the background.
10-12s:
The lens returns to a half-body of the female inner-sect disciple. She nods slightly and continues: "The Wanbao Trading Meet, held once every three years at the Holy Land's center, is about to open." As she speaks, her lip movements are natural and her pace is steady; the small cloth pouch at her waist sways lightly, and her robes and hair are gently stirred by the morning breeze.
12-14s:
The lens stays on the half-body of the female inner-sect disciple and the cave-doorway environment. She stands outside to the right, the breeze stirring her clothes, and continues: "The sect is bringing a few disciples along to see the world this time, and Senior Sister Su named you." The background stays on the cave doorway and mountain mist; do not cut to the sect mountain gate, do not show the Holy Land's center, do not show the trading meet venue.
14-15s:
The lens returns to the cave doorway. Ye Xingchen nods slightly after hearing this, his face still calm, a fleeting shift in the depth of his eyes. The last frame is fixed on: Ye Xingchen standing inside the cave doorway to the left, the female inner-sect disciple standing outside to the right, morning mountain mist drifting, Ye Xingchen's expression calm but his eyes shifting faintly.
Important requirements:
Ye Xingchen and the female inner-sect disciple must be two different people. Ye Xingchen uses Image 1; the female inner-sect disciple uses Image 2. The two must not share the same face and must not look like twins. Do not show the sect mountain gate, the Holy Land's center, or the trading meet venue. No subtitles, no on-screen text, no background music.
Reference materialPromptOutput video

Image 1

2dd2247bc3f54ffaa3cfc6700c6e0a93

Image 2

cc28634edf834ef6ada5d12fc365b7ef

Image 3

1bace15f0fe64eeca56b9acd791443e5

Video 1

Referencing the characters of Image 1 and Image 2, and referencing only the video motion of Video 1, generate a 14s fight video. The characters' appearance, hairstyle, build, and clothing follow the provided images and stay consistent throughout. The characters replicate all of the actions, postures, rhythm, and pauses from the reference video; the sequence, amplitude, speed, and timing of movements stay consistent, and the original video's music is retained. Expressions change naturally with the rhythm of the action. The reference video is used only for motion control and its character appearance is not inherited. The final image is in normal color, the characters are stable, the motion is fluid, with no subtitles and no watermarks. The characters' faces are stable, their body structure is correct, fingers and limbs are natural, with no limb interpenetration, no posture errors, and no deformed gestures; if the reference video shows obvious hand or limb deformities, especially during turns — if the hands show deformity or incoordination — the motion may be lightly adjusted.

The scene uses the environment of Image 3: a cyberpunk-style future Chinatown street, with Chinese-style gateways and neon signs, red lanterns, holographic billboards, stone lions, and wet reflective asphalt; in the distance, soaring future-city towers and flying vehicles. The two square off face-to-face on this street; the environment's building structures, sign placement, lighting mood, and color tone stay as shown in Image 3 throughout, and must not be replaced with other scenes. The Image 1 character (a black-haired girl in a dark gold-floral embroidered qipao vest, white shirt, and long braided hair) and the Image 2 character (a girl with long white hair, red combat suit, and silver armor) are the two fighters; each character's look is locked to their own reference image, with no mixing or swapping of clothing or hair color.

E-commerce advertising

Reference materialPromptOutput video

Image 1

image1

Image 2

image2

Image 1 is the main subject character; generate a video based on the content of Image 2.

Image 1

1783314905158image

Image 2

93fa9e19f11e4802b0b1f7990eb1a9dc

Video 1

Convert Video 1 into an image style: a hair dryer exploded-view diagram, detailed internal components, ring-shaped chips, an internal ring-shaped motor module, motion graphics, 4K high definition.

Gaming

Reference materialPromptOutput video

Image 1

fd448d63f3cc41ec80dc2bd2531ba969

Image 2

63633400c9d4475da347f40c696d0231
Ice-and-snow steampunk build animation, fixed high-angle overhead shot. Keep the composition, camera angle, lens height, and terrain range consistent with reference Image 1 (first frame) and Image 2 (last frame). The scene is an arctic snowfield; scattered pine trees, rocks, and distant snow textures on all sides remain stable and unchanged throughout. The entire film starts from a single furnace standing in the middle of the empty snowfield and ends with a complete steampunk village built and a train pulling into the station. The visual style is a polished game cinematic; the cold blue ambient light of the snow contrasts strongly with the orange-red hot light of the furnace. The overall motion is an orderly, clean, mechanical, building-from-nothing generation process — no chaotic explosive growth.  The scene is a flat clearing in an arctic snowfield.
 The opening shows only the furnace standing at the center of frame. The furnace is a dark-metal cylindrical structure with a chimney on top emitting white steam, and orange-red firelight glowing inside the furnace mouth. The heat of the furnace forms a warm circular halo on the snow, the halo's edge gradually transitioning into the cold blue snowfield. Pine trees and rocks are distributed in the distance and along the frame edges, their positions never changing. All subsequent buildings grow outward from the furnace as the center, ultimately forming a rectangular steampunk village.
 Shot 01, fixed high-angle establishing. The camera holds a high oblique overhead angle, wide-angle composition, with the furnace at the center of frame. As the shot begins there are no other buildings, no enclosure walls, no rails, no platform, no train. Only the furnace burns alone, its orange-red hot light illuminating the surrounding snow. White steam continuously rises from the top of the furnace, slowly drifting upward. The area of snow lit by the firelight takes on a soft warm tone, while the distant snowfield remains cold blue. The base layer of sound is cold wind, the furnace's low burning rumble, the slight vibration of the metal furnace body, and steam pressure-release hissing. This shot is relatively brief, only establishing that the furnace is the core of the entire village's generation.
Shot 02, first ring of buildings generated. The camera remains fixed, angle unchanged. A faint golden construction light pattern begins to appear in the snow around the furnace, spreading outward from the furnace base like a mechanical blueprint. The wooden cabins closest to the furnace generate first: foundations surface from beneath the snow, then walls quickly assemble, roofs drop into place, and chimneys rise up. The workshop on the other side generates more slowly: first a metal frame appears, then brick walls, gears, piping, and a small chimney. The warehouse generates quickly, with wooden beams, side walls, and the slanted roof joining almost continuously. Each independent building generates at a different speed — some snap into form rapidly, some are built layer by layer, some show the frame first before details are filled in. Do not let all buildings generate in sync; do not let them pop up simultaneously like copy and paste. The generation process must have a clear sequence and sense of layering.
Shot 03, second ring of buildings expands. Centered on the furnace, more steampunk buildings generate progressively outward. A power-generation tower slowly rises on one side, its metal supports locking into place section by section, the blue energy core at the top lighting up last. Multiple small workshops and residential cabins appear in succession at different positions, with roofs, window frames, porches, and chimneys filled in one by one. Steam pipes extend from the furnace base in all directions, connecting to each building like metal blood vessels. Some pipes lay out halfway, pause briefly, then continue on to connect to their corresponding building. The snow on the ground is pushed aside by the generating buildings, revealing wooden platforms, stone paving, and metal foundations. As buildings generate, they are accompanied by faint metal joining sounds, wood-panel landing sounds, gear-meshing sounds, and steam-jet sounds. The village gradually becomes busier, but the frame remains clearly orderly.
Shot 04, complete interior village takes shape. The camera holds the same overhead angle, with an extremely slight slow pull-back possible so that the outline of the entire rectangular village gradually comes into view. The furnace remains the center of frame and the brightest heat source. All interior buildings are essentially complete: cabins, workshops, warehouses, the power tower, steam pipes, and roads form a compact layout around the furnace. The windows of different buildings light up with warm glow in succession, some chimneys begin to smoke, and some machines begin to turn. Compacted road textures and wooden-platform edges appear on the snow. Do not change the positions of pine trees and rocks; do not let natural features be swallowed by buildings. At this point the village interior is complete, but the perimeter wall does not yet exist.
Shot 05, enclosure wall generated. The perimeter of the rectangular village begins generating an enclosure wall. The wall does not appear all at once; instead it is built in segments from the four corners and several gatehouse positions. First gatehouse foundations appear, then mixed wood-and-metal walls rise, followed by railings, support posts, and steam lamps. The wall extends along the rectangular boundary in both directions, finally closing at the main gate on the front side and the flanking gatehouses. The wall generates faster than the interior buildings, but each segment can still be seen landing in sequence. When the wall closes, the lamps on the gatehouses light up one by one, and the village's boundary becomes clearly defined. The base-layer sound adds heavy wooden-post landing sounds, metal rivet-locking sounds, and the mechanical clatter of gatehouses closing. Do not let the wall appear before the interior buildings; the wall must be generated only after all major buildings are complete.
Shot 06, rails and platform generated. After the wall is complete, two parallel dark tracks begin to appear on the snow outside the upper edge of frame. The rails extend from the distance in the upper part of frame toward the village's side; sleepers drop into the snow one by one, and steel rails then lock onto them. A platform beside the rails begins to generate: first a wooden-platform foundation appears, then the platform edge, canopy, lamp posts, and small steam signal lamps. The platform forms a clear connection with the village entrance. The snow is pressed into realistic indentations by the rails and platform, with a small amount of pushed-aside snow piled up alongside. The rails generate more linearly than the buildings, like a route being laid from the distance toward the village. Do not let the train appear early; the train must appear only after the rails and platform are fully generated.
Shot 07, train pulls in. After the rails and platform are fully formed, a steam train appears in the distance in the upper part of frame. The train travels along the rails from the upper part of frame toward the platform, with a clear direction, gradually approaching the platform. The locomotive emits thick white steam, the wheels and connecting rods move mechanically and clearly, and the body is a steampunk style combining dark metal and wooden carriages. The train gradually slows down, releasing a large cloud of steam as it approaches the platform. The platform lights come on, and the signal lamps shift from cold to warm. The train finally stops smoothly beside the platform, the wheels stop turning, and steam spreads to either side. The furnace still burns at the center of the village; all buildings, the enclosure wall, rails, platform, and train form the complete final frame.
 The film ends on a stable overhead composition of the complete village and the stopped train.
Throughout the film, the physical effects must be clearly visible: the furnace's orange-red hot light illuminating the snow, white steam rising, the warm-cold light transition on the snowfield, building foundations surfacing from the snow, wood panels and metal frames joining, chimneys rising one by one, steam pipes extending segment by segment, window warm lights lighting up one by one, the enclosure wall closing in segments, rail sleepers dropping into place one by one, pushed-aside snow forming piled edges, the train's wheels turning, and steam jetting out and diffusing in the cold air.  In the first section the lighting uses the furnace as the sole warm light source, with the surrounding snowfield maintaining cold blue ambient light. After the buildings are generated, windows, steam lamps, and signal lamps gradually add new warm point light sources. Once the village is complete, the central furnace light, building window light, wall lamps, and platform lamps together form a warm core. When the train pulls in, the locomotive's headlight sweeps a beam of warm light across the snow, and the steam is lit by the light. Maintain a clear overhead game-cinematic quality throughout: high detail, sharp building edges, clean snow textures.  Do not change the camera angle or composition. Do not let pine trees, rocks, or distant natural terrain drift. Do not let all buildings generate simultaneously. Do not let the enclosure wall appear before the interior buildings. Do not let the rails and platform appear before the enclosure wall. Do not let the train appear before the rails and platform. Do not let buildings interpenetrate, overlap, or float. Do not let a building disappear after being generated or have its position jump. Do not let the train derail, pass through buildings, or stop outside the platform. Do not let the frame flicker, drop frames, shake, or switch viewpoints mid-way. Do not show close-ups of people. Do not show text, watermarks, interface buttons, or game UI.

Image 1

storyboard reference image, five-level material version
Bright, fresh Pixar-style 3D animated film quality, smooth glossy rounded shapes, saturated translucent colors, high-key lighting, blue sky and white clouds, sunlight piercing through the jungle canopy casting dappled light, lush emerald-green leaves, warm-yellow window light, cheerful and bright as a fairy-tale picture book. No darkness, no gloom, no grime, no horror, no photorealistic style. Vertical 9:16 frame.
The protagonist is a spirited cartoon girl: dark-brown high ponytail, a red plaid bandana tied around her forehead, large bright eyes, a slate-blue multi-pocket utility vest over a white short T-shirt, khaki rolled-cuff cargo shorts, knee pads, fingerless gloves, and a tool belt at her waist. The toxic fog is a softly glowing mint-green heavy gas that sinks and slowly churns and flows along the forest floor; the zombies are round-faced chubby figures with mint-green skin, buck teeth, and short outstretched arms — an adorable cartoon look, not scary at all.
This is a 30-second long take shot in one continuous shot from beginning to end. Absolutely no hard cuts, no jump cuts, no editing, no scene changes, no transition effects. The camera does one thing the entire time: it smoothly lifts upward while simultaneously pulling back, rising from ground level all the way to four hundred meters altitude, the motion never stopping, never changing position. All construction and transformations must be performed as visible animation on screen — structures must physically grow, be hoisted, laid out, slide out, unfold, and assemble piece by piece; no jumping directly to the finished state.
[Overall sound] Throughout the film, a young female voice narrates in first-person Mandarin Chinese, standard and clear pronunciation, the voice bright and clean with a slightly husky firmness, as if recounting her own lived experience. The first two lines are slightly fast, with breathless urgency; the middle lines turn warm and grounded; the final two lines are confident and rising, gaining conviction as they go. The narration volume always sits above the ambient sound. Also include a progressively building adventurous brass-and-drums score, with a heavy downbeat at the moment of each treehouse upgrade, peaking in the final four seconds and not receding. Do not show any text, subtitles, numbers, interface elements, or watermarks on screen.
Seconds 0 to 3: Looking up through gaps in the jungle leaves at the blue sky, a damaged white passenger plane dragging billowing white smoke crosses the frame diagonally with a slight roll, debris falling through the canopy, foreground leaves trembling as it passes. Sound effects: the piercing whine of the engine, metal tearing, and leaves rustling. From 0.3s the narration says: "A plane crash dropped me into this toxic-fog forest."
Seconds 3 to 6: The camera tilts down; the girl sprints at full speed along a sunlit-dappled forest path, glancing back in panic, her ponytail and bandana tails whipping, chased by a crowd of adorable mint-green zombies with short arms stumbling after her — one trips and falls flat, the smoking plane wreckage visible in the distance. Sound effects: hurried footsteps, panting, zombie groans, and branches snapping. From 3.2s the narration says: "One breath of the toxic fog and you turn into a zombie; I could only run upward."
Seconds 6 to 8.5: The girl grabs a hanging vine, kicks off the bark protrusions, and climbs a giant tree covered in moss and vines; the vine swings as it is pulled, moss debris drifting down, the crowd of zombies below reaching upward futilely in the mint-green toxic fog, gradually shrinking and sinking into the fog as the camera rises and pulls back. Sound effects: climbing friction, the creaking of vines under strain, and the zombies' hissing fading below. From 6.2s the narration says: "This giant tree became my only home."
Seconds 8.5 to 13, first build: She sits on an empty branch fork and catches her breath; then a section of white airplane fuselage wreckage is hoisted up by a creaking rope-and-pulley system, swaying and wobbling as it slots into the fork, its muffled landing shaking loose a shower of leaves; wooden planks are passed up one by one and automatically lay out from the center outward in front of the fuselage to form a platform; next a hammock stretches between two branches, a campfire poofs alight in an iron basin, the first lantern is hung at the platform edge glowing warm yellow, and a wooden ladder unfolds down the trunk; a fluffy small dog pokes its head out of the fuselage hatch, runs onto the platform, and wags its tail. Sound effects: the winch creaking, the thud of planks landing, hammering, the campfire poofing alight, and the dog barking. From 8.8s the narration says: "Salvaged a piece of fuselage wreckage — this was my first home."
Seconds 13 to 17.5, first upgrade, wrap-around growth: She works the winch to hoist up a bundle of logs; the logs stack themselves horizontally one by one, piling around the trunk and the old platform into a sturdy log wall, each log landing shaking loose wood shavings; the airplane fuselage does not disappear — it is wrapped inside the new wall, becoming the entire left wall of the cabin, with its porthole still clearly visible; wooden shingles unfold from the ridge outward across the roof, a stone chimney grows out of the roof and immediately emits white smoke; window frames snap in, warm-yellow light glows inside, balcony railings enclose the space, potted plants are set out one by one, a clothesline is strung with colorful clothing, and a wooden staircase spirals downward around the trunk; the air is full of flying wood shavings and backlit dust. Sound effects: the continuous muffled thuds of logs landing, the clatter of shingles flipping into place, stone-stacking sounds, and window-sash creaking. From 13.2s the narration says: "Fighting zombies, scavenging supplies, earned my first real house."
Seconds 17.5 to 22, second upgrade, poured-and-formed: Wooden formwork panels stand up one by one and close around the log cabin to form a tall cylindrical mold; gray concrete is poured in from the top, the slurry level visibly climbing to fill the mold, while the formwork is simultaneously extended upward by two stories to raise the tower; then the formwork bursts outward and clatters apart all at once, revealing a smooth light-gray concrete round tower, dust billowing; metal pipes burst out through the walls and coil around the tower, a circular airtight hatch screws itself into the wall, the wheel-lock tightens, and a searchlight extends from the tower top and switches on, its beam sweeping across the fog below; a concrete staircase unfolds along the side step by step, green vines rapidly climb the tower from the base, and the small dog runs out of the hatch and lies on the steps. Sound effects: formwork landing, the gushing of concrete being poured, the explosive burst of demolding, the metal friction of pipes emerging, the click of the wheel-lock snapping shut, and the hum of the searchlight starting up. From 17.7s the narration says: "A tower poured from concrete — the toxic fog can't get in anymore."
Seconds 22 to 26, third upgrade, steel-stacked armoring: Steel beams are extended upward section by section from the top of the concrete tower like auto-assembling scaffolding, bolts self-threading into each joint with a flash; multiple steel decks unfold layer by layer from the central axis outward, railings enclosing as they go, and steel plates slide in from the outside to clad the tower like armor, welding arcs flickering along the seams; a solar-panel array unfolds and turns toward the sun, a rooftop wind turbine's blades are installed and spin from slow to fast, and a cargo cage is hung on a steel cable and begins ascending and descending the tower; a ring of electrified-grid posts rises at the tower base, bright blue-white arcs lighting up one after another with a crackle, stringing into a sizzling halo that carves a clear indentation into the mint-green toxic fog below, and a few adorable zombies harmlessly topple backward; the warm-yellow lights in the tower's windows light up layer by layer from bottom to top, and the girl rides the cargo cage up to the top floor and waves back. Sound effects: the clanging of steel beams joining, pneumatic bolt-tightening, the sizzle of welding, the whoosh of the wind turbine speeding up, the hum of the electrified grid, and the crackle of arcs. From 22.2s the narration says: "Power it up, mount the guns — this forest is mine."
Seconds 26 to 30, final upgrade, energy reconstruction: A blue holographic grating sweeps upward from the tower base across the entire steel fortress; where it passes, the structure resolves into glowing blue wireframes; the wireframes disintegrate, the steel components shatter into hundreds of cubes that float upward and rotate in mid-air, briefly revealing the giant-tree trunk still standing at the core; the cubes settle and recombine into a sleek white-blue saucer-shaped platform, unfolding layer by layer from the central axis and spreading outward, five levels in total and smaller toward the top, each level emitting a ring of blue light at its edge as it unfolds; a bright blue energy beam lights up from the tree root straight through to the tower tip, a transparent dome closes from the edge inward over the roof garden, and a waterfall-like curtain of water pours from the platform's edge; finally a swarm of small drones takes off and circles the tower, greenery and warm light showing through inside the dome, and on the hazy golden horizon two or three more identical megastructures glow with lights; the camera pulls back dramatically to become a high-altitude wide shot of the entire magnificent megastructure. Sound effects: the scan's electronic tone, the hum of the structure disintegrating, the mechanical sound of the saucer layers unfolding, the low-frequency rumble of energy infusion, and the buzzing of the drone swarm; the music peaks. From 26.2s the narration says: "From a piece of wreckage, to a city — how high can you build?"
Throughout the film, the trunk and vines of that giant tree are always visible; the buildings only grow winding around it. At every stage it must be readable as a treehouse, not a tower.

Voiceover version:

No-voiceover version:

Creative use cases

Text-to-video

PromptOutput video
[Style/Lighting/Atmosphere] 3D Pixar/Disney animated film quality, hyper-realistic plush-material rendering, fluffy soft fur detail, soft warm key light + slight side backlight, delicate highlight-to-shadow transitions, clean translucent frame, minimalist black background to make the subject pop, high-saturation warm color palette, healing cozy atmosphere, comedic contrast-cute, adorable soft-squishy character expressions, natural emotional progression, 8K ultra-clear, cinematic detail, cartoon render, The Secret Life of Pets / Zootopia style, no text or watermarks, pure visual. [00:00-00:02 medium close-up] A short-legged corgi sits on a wooden floor, its round rump planted, tail wagging frantically, grinning happily, face full of excited anticipation, staring at a rolling colorful ball of yarn, warm light on its fluffy orange-and-white fur, soft highlights at the fur's edges. [00:02-00:04 close-up] The corgi leans in to grab the yarn ball in its mouth; the ball suddenly veers and slips away; the corgi tilts its head and squints, face full of confused puzzlement, the yarn ball's reflection shining in its bright black eyes, soft reflections forming on the yarn ball's surface. [00:04-00:07 fast cutting] The corgi repeatedly lunges, bites, and paws, failing again and again; the yarn ball rolls back and forth, constantly dodging; the corgi shifts from puzzled to agitated and annoyed, its tail going from wagging to stiffly swaying; the warm lighting sways slightly with each lunge, heightening the comedic rhythm. [00:07-00:10 tempo increases] The corgi rises on its hind paws and pounces, chasing back and forth, short legs bobbing as it runs, its round rump swaying along, the yarn ball bouncing elastically and veering in exaggerated escapes, the lighting tracing the corgi's round body silhouette. [00:10-00:12 slow motion] The corgi's front paws pin the yarn ball firmly, but the ball instantly slips free and escapes; the corgi lunges and misses, flops on the ground, head resting on its paws, face full of frustrated disappointment; in slow motion the fluffy fur texture and the soft light feel even cozier. [00:12-00:15 climax] The corgi bounces in place and paws the ground, the yarn ball bounces up and smacks its face then slides off; the corgi first goes blank with shock, eyes wide, then its face falls, ears drooping, finally freezing in a breakdown as warm highlights hit its pitiful little face — funny and heartwarming at once.
Core: cinematic hyper-realism, first-person LEGO-brick perspective, one continuous shot gliding smoothly through, hardcore adventure with a cold hard-edged texture; no dialogue, no BGM, only high-fidelity ambient sound; 15 seconds of high-speed thrilling danger-evasion throughout, amplifying the LEGO plastic texture and the sense of scale contrast.
Visuals: LEGO-brick perspective bursting into a giant human home scene; floor = wilderness, tabletop = wasteland, books = cliffs, water cup = skyscraper tower, keyboard = machinery cluster; hands and footsteps bring intense oppressive presence, no cartoon feel, faithful to home-scene logic.
Shot list (15 seconds)
0-2s: miniature low-altitude POV, launched from the edge of the floor/tabletop, skimming the ground at full speed, ultra-wide angle amplifying scale contrast, quickly pulling the viewer into the micro-adventure.
2-5s: high-speed traversal, brushing past LEGO parts, books, pens, keyboards, water cups, dodging the shadow of a hand, skimming the ground / diving / squeezing through gaps, extreme obstacle avoidance.
5-8s: high-risk threading, sliding toward the sofa, threading through chair legs, rushing the book-page cliff, dodging the water cup / human footsteps, tempo increasing, danger maxed out.
8-11s: breaking into the human activity zone, dodging typing, page-turning, and drinking motions; keyboard roaring, page-turning like a gale, extreme danger-evasion threading.
11-13s: bursting out of danger, entering the half-assembled / parts zone of LEGO, atmosphere briefly easing, clarifying the LEGO perspective.
13-15s: sharp pull-up and freeze-frame on the finished LEGO build, leaving giant home-scene afterimages in the background; a sharp, memorable finish.
Sound: no BGM; opening wind-shear / plastic-friction sounds; mid-section keyboard / paper / footstep / vibration sounds; latter half plastic snap-fit / faint wind sounds, amplifying the scale contrast.

LOGO growth animation

Reference materialPromptOutput video

Image 1

Asset 2

Image 2

Asset 3

Video 1

First frame: Use Image 1 as the first frame at the start of the video.

Last frame: Use Image 2 as the last frame at the end of the video.

Reproduce the logo-appearance animation effect and the text-appearance dynamic effect from Video 1, and generate a professional MG-animation style (Motion Graphics, dynamic graphic design). The overall visual presentation should carry a "tech feel, beat-synced rhythm, clean and crisp" tone.

Image 1

Asset 3
Image 1 is the core visual anchor and material reference. Against a deep pure-black background, emerald-green (matching the Logo color value) digital light particles rush from all sides of the frame toward the center at extreme speed. The particles take on a pixelated or low-poly style, trailing ultra-fine green fluorescent tails, drawing tech-charged trajectories across the black space.
The points grow denser, weaving at the center of frame into a flowing energy vortex that faintly refracts a cold glass- or metal-like sheen. Suspended green motes drift slowly through the space, simulating the visual of data flowing.
Phase 1 (0-3 seconds): Just left of center, scattered cyan-green points spiral inward from all directions, speed slowing then quickening, their tails bending in zigzag patterns in the direction of motion (echoing the Logo's geometric feel), the light trails lingering for about 0.3 seconds before dissipating.
Phase 2 (3-5 seconds): The particles converge faster, densely forming a high-luminance orb about one-third the width of the frame; at the orb's center the silhouette of the QwenWork Logo gradually emerges, a holographic-projection-like scan sheen appearing inside, with strong green outer glow at the edges, as if the system is booting up.
Camera: slow push-in, the rig advances at a steady pace from a medium shot toward the central orb, keeping the composition symmetric and emphasizing focus.
Sound: crisp electronic startup tone + the subtle electric-current sound of data flowing, closing with a single deep, forceful boom.
Duration: 5 seconds, ratio: 16:9.

Dance motion replication

Reference materialPromptOutput video

Image 1

female lead identity version

Image 2

four backup dancers identity version

Video 1

Reference Video 1's camera movement, shot size, shot rhythm, and choreography. Use Image 1 as the center-position female lead, replacing the male lead in the original video; follow the original video exactly, with the character entering from the right side of frame. Keep the character's appearance, hairstyle, body type, and temperament highly consistent, and faithfully recreate all of the male lead's dance moves and rhythm from the original video; use Image 2 as the surrounding backup dancers, matching the original video's backup-dancer positions, formation changes, and synchronized moves. The overall look is hyper-realistic cinematic quality, fusing the camera language of a Hollywood song-and-dance action blockbuster; the moves are silky smooth, the camera movement stable and fluid, the rhythm precisely beat-synced.

Image 1

sweet girl in bedroom

Video 1

Reference the character from Image 1; use Video 1 only as a motion reference. Generate a one-take, real-phone-shot-quality dance video in a pure-white shadowless studio: bright even soft light, clean background, the camera essentially fixed, with no transitions and no cuts. The character's face, hairstyle, body type, and clothing should follow the provided image and remain consistent throughout. The character replicates all the moves, poses, rhythm, and pauses from the reference video; the order, amplitude, speed, and timing of the moves stay consistent, and the original video's music is preserved unchanged; expressions change naturally with the rhythm of the motion. The reference video is used only for motion control; do not inherit its character appearance. The final frame should be normal color, the character stable, the motion fluid, with no subtitles and no watermarks. The character's face is stable, the body structure correct, the fingers and limbs natural, with no limb interpenetration, no pose glitches, and no deformed gestures; if the reference video shows obvious hand/limb structural deformation — especially during turns, with deformed or uncoordinated hands — the motion may be lightly adjusted.