Skip to content
AI Video Workflows3 min read

Wan 2.2 ComfyUI Workflow: Setup and First Video

Set up a Wan 2.2 image-to-video workflow in ComfyUI. Match the model variant, load the required files and test one approved creator image.

AI-generated illustration: A miniature city is filmed beside a video-editing workstation.

A Wan 2.2 ComfyUI workflow is a model-specific video graph. The useful first milestone is a short, coherent clip from one approved image—not the largest resolution or longest sequence your settings allow.

Choose the right Wan 2.2 variant

The official Wan 2.2 repository distinguishes T2V-A14B for text-to-video, I2V-A14B for image-to-video and TI2V-5B for text-and-image-to-video. They have different implementations and resource requirements. A filename containing “Wan 2.2” is not enough to establish compatibility.

Starting pointWhat to look for
A written scene onlyA text-to-video example for the selected variant
An approved portrait or product sceneAn image-to-video example with an image input
A smaller local testThe documented smaller-variant workflow and its actual requirements

Load the official ComfyUI example

Use the ComfyUI Wan 2.2 tutorial and its matching template. The required components can include diffusion-model weights, a text encoder and a VAE, placed in their documented directories. Use the precise files listed for the chosen graph and update through the supported installation procedure when required nodes are unavailable.

Do not mix a native graph with a community wrapper’s model instructions. Both can be valid approaches, but their loaders, supported formats and settings may differ. Follow one complete maintained example before adapting it.

Use a clean starting frame

For a creator clip, choose a face that is large enough to inspect, a plausible pose and a simple background. Fix visible anatomy or product errors before animation. The model may carry a small flaw through every frame, turning a simple image repair into a difficult video problem.

Match the composition to the intended motion. If you want the character to turn slightly, leave room around the head and shoulders. If you want a small gesture, begin with clearly visible hands that are not tangled with an object.

Run one controlled test

  1. Keep the official graph’s initial settings and record the model variant.
  2. Load the approved starting image.
  3. Write a simple motion instruction: one action and one camera behavior.
  4. Generate one clip and save the workflow with the result.
  5. Watch the full clip, including its final frames.
  6. Change one variable before the next attempt.

An original first test could be: “The adult creator makes a small relaxed head turn toward the window and settles. The camera remains fixed; the room and jacket stay unchanged.” This is a diagnostic prompt to adapt, not a claim of a tested result.

Memory requirements depend on the implementation

Do not copy a VRAM claim from an unrelated setup. The official repository and ComfyUI tutorial describe different execution configurations, including offloading choices. Requirements also depend on the variant, precision, dimensions, frame count and other running workloads. Read the requirement attached to the exact graph you are using.

If a small test fails, preserve the log. Distinguish a missing model, a loader mismatch and an out-of-memory error before changing settings. The connection and crash checklist can help separate these cases.

Plan longer videos as an edit

A longer campaign video does not have to be one long generation. Plan a sequence of independently reviewed shots: creator introduction, product detail, demonstration and closing card. This gives you specific places to replace a failed segment without rebuilding the whole edit.

Check continuity between clips: identity, clothing, product shape, light and screen direction. Read the Wan 2.2 prompt guide for shot-level examples.

If your priority is a hosted character-and-content workflow, create your AI influencer in Clout and build a small photo and video set. That is a separate workflow; do not assume Clout imports the local Wan graph or exposes identical controls.

Questions, answered

Frequently asked questions

Which Wan 2.2 model should I use for an image input?

Use the image-to-video or text-and-image-to-video variant with its matching supported template. The exact choice depends on your resources and requirements.

Can I use any Wan model file in the same graph?

No. Match the variant, format, loader and companion files to the documented workflow.

Does Wan 2.2 have one universal VRAM requirement?

No. Requirements vary with the implementation, model variant, precision, offloading and generation settings.

Keep building

View all guides