Skip to content
AI Model Guides10 min read

Seedance R2V guide for consistent creator videos

Learn Seedance 2.0 R2V with reference roles, three creator prompts, downloadable image assets and practical checks for identity, props and motion

AI-generated editorial illustration: Original fictional creator with a dark bob and sage overshirt beside a coral tumbler in a daylight kitchen used as reference artwork

Seedance 2.0 R2V means reference-to-video: you direct a new video using reference assets as well as a written prompt. A creator photo can define the person, a product image can define the prop and a short reference clip can guide camera movement. The useful step is assigning each input a clear role, rather than uploading several files and hoping the model knows which details matter.

This guide focuses on the R2V workflow for creator scenes and product moments. It includes two original downloadable image references, three ready-to-adapt prompts and a blank review sheet. The images are reference artwork, not frames from a Seedance generation. For versions, access and the wider model family, read the Seedance overview.

Understand R2V versus image to video

In a conventional image-to-video job, one image often acts as the starting composition you want to animate. In a reference-to-video job, your assets can supply separate ingredients for a newly directed scene. The distinction is about the operation and controls exposed by your chosen provider, not simply how many images you upload.

ByteDance's official Seedance announcement describes combined text, image, video and audio inputs. Its R2V examples explicitly assign different references to a character, scene, props and shooting script. These are documented capabilities; they are not a guarantee that every generated shot will preserve every detail.

WorkflowUseful starting materialGood first job
Text to videoA written scene briefExplore a new setting or visual idea
Image to videoAn approved compositionAnimate one simple action in that composition
Reference to videoAssets with explicit identity, prop or motion rolesCreate a new directed scene around those references
Video editing or extensionA clip you want to change or continueRevise a specific part using the supported editing operation

Choose image-to-video if your main requirement is to bring one finished first frame to life. Choose R2V when the brief needs a reference person in a new composition or a combination of identity, prop and movement. A reference clip is not automatically an instruction to reproduce every object or person inside it.

Build a small reference pack with clear roles

Start with the fewest assets that explain the scene. More references can introduce competing outfits, lighting and faces. For a simple creator product moment, begin with one identity reference and one prop reference. Add a motion clip only if it communicates something your written camera instruction does not explain well.

Asset roleWhat to preserveWhat to exclude from that role
IdentityFace, hair and approved wardrobeUnrelated people or conflicting outfits
Product or propShape, color and visible constructionInvented functionality or unverified claims
SceneLayout, lighting and atmosphereAn unwanted background character
MotionOne action or camera trajectoryThe reference clip's identity unless requested
AudioThe intended sound roleUnspecified copying of every sound in the file

Use an original fictional character or a person whose likeness you are permitted to use. Keep the asset names meaningful, such as creator-reference and prop-reference. Review the files at full resolution before uploading them: the prompt cannot reliably rescue an obscured face or a product image with a missing lid.

Keep a written role map next to the files. If a provider labels uploads by their order, reordering them can change what a token points to. Check the displayed labels after removing or replacing an asset. The consistent creator workflow helps you establish the identity baseline before directing motion.

Download the original creator and prop references

Our example uses a fictional creator with a dark bob and sage overshirt at a kitchen table. The coral tumbler with a cream lid is a fictional prop. We created the scene as original artwork, then made a separate reference-based product image. The prop edit preserves the broad concept, but it is not a pixel-identical extraction or evidence of a real product's dimensions.

Original fictional creator with a short dark bob and sage overshirt beside a coral tumbler at a daylight kitchen table
Original identity and scene reference artwork rather than a generated video result

Download the creator reference. For the examples below, assign it to the identity role. The kitchen can also guide the first scene, but it should not force every later prompt to use that location.

Separate original reference artwork of a matte coral cylindrical tumbler with a cream lid on a warm neutral studio surface
Separate fictional prop reference for shape and color review

Download the prop reference. Inspect the cream lid, cylindrical body and rounded base before using it. For an actual ecommerce campaign, substitute your approved product photography and verify the object's shape throughout the finished clip.

Write a prompt that assigns references explicitly

A practical R2V prompt has five parts: asset roles, scene, action, camera and constraints. Write the role assignment first, then specify one observable action. “Make an amazing ad” leaves the production decisions unresolved. “Lift the coral tumbler once and return it to the table” gives you a concrete event to check.

Reference-label syntax depends on the interface. The fal Seedance reference-to-video schema uses tokens such as @Image1, @Video1 and @Audio1. Use the labels your provider actually displays. Do not assume a token copied from another interface still points to the intended file.

Use @Image1 for the creator's identity, short dark bob and sage overshirt. Use @Image2 for the coral tumbler's shape and cream lid. One continuous medium shot at a warm kitchen table. The creator looks toward camera, lifts the tumbler once to chest height, smiles naturally and returns it upright to the table. Camera remains still. Preserve the face, wardrobe and prop colors. No new people, no added logos and no scene change.

This is an original prompt to test, not a report of a successful Seedance render. The two downloadable images supply its image references. No motion or audio clip is required for this first creative brief. The camera can remain still while you evaluate identity and hand contact.

Three prompts for different creator use cases

Creator product introduction is the simplest starting point. Use the kitchen prompt above for one lift-and-return action. Keep the prop visible and avoid opening the lid, pouring liquid and speaking at the same time. Those additional tasks can be tested after the basic action passes review.

Motion-led campaign variation adds a permitted camera reference. Only use this prompt after supplying your own appropriate video asset and checking its displayed label.

Use @Image1 for the same creator and wardrobe. Use @Image2 for the coral tumbler. Use @Video1 only for its slow forward camera movement, not its people, clothing or location. Place the creator at a quiet cream cafe table in daylight. She rests one hand beside the tumbler and glances from it toward camera. One continuous shot with a gentle forward move. Keep the creator and prop recognizable and keep the cream lid attached.

Short story reaction changes the narrative without adding a complicated physical task. A small expression can create a clearer test than a large action sequence.

Use @Image1 for the creator's identity, hair and sage overshirt. Use @Image2 for the coral tumbler and cream lid. One medium shot at the kitchen table. The creator notices that the tumbler is already beside her, briefly raises one eyebrow, then gives a small amused smile toward camera. Keep both hands resting naturally on the table. Locked camera, soft daylight, no extra objects, no transformation and no added text.

Download all three prompts and the reference role map. Replace the fictional scene details with your own brief. If you add dialogue, write a short line and review the spoken export; do not infer exact voice identity from the still reference. Add campaign captions in your finishing workflow so spelling and placement remain controllable.

Check the provider settings before you generate

The fal schema checked for this guide lists up to nine image, three video and three audio references, with no more than twelve files total. It requires at least one image or video reference. Output duration options run from four to fifteen seconds or auto. These are provider-specific documented limits, not universal controls across every Seedance interface.

For a first test, choose a short duration and one aspect ratio appropriate to your destination. Record the exact endpoint or model label, duration, resolution and audio setting. If the interface does not expose a control mentioned in another provider's documentation, do not hide that difference behind the general Seedance name.

Our example starts with two still references and one continuous shot. It leaves video and audio inputs out until they have a clear role. Uploading empty fields or naming a nonexistent reference in the prompt does not create the missing input.

Confirm the current charge in the provider before submitting. Budget for accepted clips rather than assuming every attempt is publishable. This article does not invent a universal free allowance, subscription price or native Seedance endpoint inside Clout.

Fix an ignored reference or drifting creator

ProblemInspect firstUseful next revision
The wrong person appearsIdentity label and competing facesUse one clean creator reference and name its role first
The product changesProp visibility and hand occlusionReduce handling and keep one side clearly visible
The motion clip dominates everythingWhich parts of the clip are requestedLimit the reference role to camera movement
The shot adds an unwanted cutConflicting shot instructionsRequest one continuous shot and fewer events
The hand merges with the lidThe contact moment and action countKeep the lid attached and test a simpler lift
The expression changes unpredictablyExtreme turns or facial actionsUse a smaller expression and restrained camera move

Change one major variable at a time. If you replace the creator photo, change the location and add dialogue in one revision, you lose a clear explanation of what helped. Save the accepted references with the corresponding prompt and settings so the next shot starts from an actual reviewed baseline.

A repeated failure may reflect an unsupported input, a provider constraint or a difficult brief. Simplification is a useful diagnostic, but it is not a guarantee. Keep the original and check the downloaded result rather than relying only on a small preview.

Review the finished sequence and build your next scene

Inspect the opening frame, the action midpoint, the ending and uninterrupted playback. In our kitchen example, the specific checks are the creator's face and bob, the sage wardrobe, the coral body and cream lid, the grip during the lift and the return to the table. Look for changes during motion, not only at the first frame.

Download the blank R2V review sheet. Its observation and decision fields are empty. Record what happened in your own render, such as a changed lid or an unrequested cut. No quality score or measured success rate has been prefilled.

Clout can help you build an original recurring creator and create reference photos and video scenes around that identity. Establish the character you want your audience to recognize, then direct each scene with a concrete action and a consistent visual brief. R2V is one production method within that larger creator workflow.

Continue with the camera movement prompt guide for motion direction, the product photography guide for reference review and the image-to-video guide when animating a finished composition is the better fit.

Questions, answered

Frequently asked questions

What does Seedance R2V mean

R2V means reference-to-video. It uses supplied assets and a written prompt to guide a newly directed scene, with explicit roles for identity, props, setting or movement.

How is R2V different from image to video

Image-to-video often animates an approved starting composition. Reference-to-video can use assets as separate ingredients for a new scene. The available operations depend on the provider.

How do I reference an image in a Seedance prompt

Use the label shown by your chosen interface and explain its role. The fal schema uses tokens such as @Image1; check the order again when you replace or reorder uploads.

Does every R2V workflow require a reference video

No. The fal operation documented here requires an image or video reference, so a workflow can begin with still images. Add a motion clip only when it serves the brief.

Can R2V guarantee the same face and product

No reference guarantees an exact result. Inspect identity, product shape, hand contact and motion throughout the export, then revise the smallest relevant part of the brief.

Were the downloadable images generated by Seedance

No. They are original fictional reference artwork made with the built-in image-generation tool. The prompts are creative briefs to test, not a completed Seedance video benchmark.

Is Seedance R2V free

Access and billing depend on the provider and account offer. Confirm the current operation, allowance and export conditions rather than assuming a universal free plan.

Keep building

View all guides