Updated: September 3, 2026 · Docs · Video AI · Clothoff AI Editorial Team

What is reference-to-video?

A generation mode that keeps one subject consistent across new scenes using several reference images — how it differs from image-to-video and why identity locking raises the consent bar.

Technical background for readers of tool reviews; legal references are general information, not legal advice.

Reference-to-video is a generative-AI mode in which one or more reference images of a subject — a person, an outfit, an object or a place — guide the model so that the subject stays recognizable in a new scene described by the prompt. Unlike image-to-video, the references are not the first frame; they are identity constraints on every frame.

Reference-based video modes at a glance

Checked against the linked sources on September 3, 2026; no editor scores here (how we rate).

TermWhat it doesWhere usedLimits
Reference-to-video (Vidu)Accepts 1–7 images of subjects and renders them into a new sceneVidu Q3, Q2, Q1, 2.0 via API; Q3 clips 3–16 s at 540p–1080pUp to 7 subjects; identity weakens with few references
Elements (Kling)Composite assets from several angles; up to 3 bound to start or end framesKling 3.0; up to 7 reference characters in videoElements must appear in the reference frames
Ingredients to Video (Veo)1–3 reference images guide generationVeo 3.1; 8-second clips, 1080p or 4K, SynthIDFixed length; hosted only, content policies apply
Image-to-video (I2V)Animates one image as the first frameWan2.2 I2V-A14B, Veo, Vidu, Kling; undress appsSubject drifts when the camera moves
Face swapReplaces a face in existing footageStandalone tools; paid add-on in PornworksNeeds existing video; edges and lighting mismatch

How is reference-to-video different from image-to-video?

In image-to-video the picture is the opening frame and the subject may drift as the camera moves. In reference-to-video the images act as an identity embedding: the model encodes the subject’s face, body or shape and conditions every frame on it while the prompt decides scene, camera and action. The same person can appear in a room that was in no input.

Vidu’s API documentation states that the endpoint “accepts 1 to 7 images,” with image and text subjects capped at 7 in total. Kling calls the equivalent asset an Element — “a composite asset” built from multiple angles — and binds up to 3 elements to a start frame or start-and-end pair in the 3.0 model.

Which models offer it?

Vidu (Shengshu) exposes reference-to-video across its Q3, Q2, Q1 and 2.0 models; Q3 adds audio and 3–16-second clips at 540p to 1080p. Kling 3.0 supports up to 7 reference characters in video. Veo 3.1 calls the feature Ingredients to Video, uses 1–3 reference images and marks output with SynthID. Model entries: Vidu, Kling, Veo, Wan.

How do we treat it in reviews?

No undress app in the current ranking exposes a multi-reference video mode; the video features tested in Ainudez, Deep Undress and Pornworks are single-image I2V, and Pornworks adds face swap on existing footage. When a reviewed tool adds reference-to-video, the editorial team tests it with synthetic or consenting-adult inputs only, records whether it checks consent for uploaded references, and notes output marking — items feeding the privacy score in how we rate. Victims of an identity-locked clip can hash it with StopNCII.org or report it here through the NCII form. This page is general information, not legal advice.

Questions about Reference-to-video

Frequently asked questions

How many reference images does reference-to-video need?

Vidu accepts 1 to 7 images per request across its Q3, Q2, Q1 and 2.0 models. Kling 3.0 binds up to 3 elements to a start or end frame and supports up to 7 reference characters in video. Veo 3.1 uses 1 to 3 reference images. More angles generally give a steadier identity across the clip.

Is reference-to-video the same as face swap?

No. Face swap edits an existing video by replacing one face. Reference-to-video generates a completely new clip in which the referenced subject appears, including body, clothing and motion. The second is harder to detect because no original footage exists for comparison.

Do undress apps offer reference-to-video?

Not in the tools tested so far. Ainudez, Deep Undress and Pornworks provide single-image image-to-video, and Pornworks adds face swap on the Ultimate plan. Wavespeed lists general video models via API. The editorial team re-tests video features at each review update and records any change in the review.

Can a reference-to-video clip be identified as synthetic?

Veo output carries SynthID, and hosted services increasingly attach C2PA metadata. Open-weight pipelines leave no mandatory mark, so detection relies on forensic classifiers or platform hashing. From August 2, 2026 the EU AI Act requires machine-readable marking of synthetic video by providers serving EU users.

Is it legal to put a real person into a generated video?

With that person’s consent, yes. Without it, an intimate clip is a “digital forgery” under the TAKE IT DOWN Act, with up to 2 years in prison for adult subjects and 3 for minors, and covered platforms must remove it within 48 hours. Non-intimate uses can breach publicity or state deepfake laws.

Why does this site care about identity locking?

Because the privacy score in every review weighs how a tool handles other people’s likenesses. A mode built to preserve a specific person’s identity across scenes needs a consent check on uploaded references and output marking. Tools that skip both lose points, and reviews say so.