What is reference-to-video?
A generation mode that keeps one subject consistent across new scenes using several reference images — how it differs from image-to-video and why identity locking raises the consent bar.
Technical background for readers of tool reviews; legal references are general information, not legal advice.
Reference-to-video is a generative-AI mode in which one or more reference images of a subject — a person, an outfit, an object or a place — guide the model so that the subject stays recognizable in a new scene described by the prompt. Unlike image-to-video, the references are not the first frame; they are identity constraints on every frame.
Reference-based video modes at a glance
Checked against the linked sources on September 3, 2026; no editor scores here (how we rate).
| Term | What it does | Where used | Limits |
|---|---|---|---|
| Reference-to-video (Vidu) | Accepts 1–7 images of subjects and renders them into a new scene | Vidu Q3, Q2, Q1, 2.0 via API; Q3 clips 3–16 s at 540p–1080p | Up to 7 subjects; identity weakens with few references |
| Elements (Kling) | Composite assets from several angles; up to 3 bound to start or end frames | Kling 3.0; up to 7 reference characters in video | Elements must appear in the reference frames |
| Ingredients to Video (Veo) | 1–3 reference images guide generation | Veo 3.1; 8-second clips, 1080p or 4K, SynthID | Fixed length; hosted only, content policies apply |
| Image-to-video (I2V) | Animates one image as the first frame | Wan2.2 I2V-A14B, Veo, Vidu, Kling; undress apps | Subject drifts when the camera moves |
| Face swap | Replaces a face in existing footage | Standalone tools; paid add-on in Pornworks | Needs existing video; edges and lighting mismatch |
How is reference-to-video different from image-to-video?
In image-to-video the picture is the opening frame and the subject may drift as the camera moves. In reference-to-video the images act as an identity embedding: the model encodes the subject’s face, body or shape and conditions every frame on it while the prompt decides scene, camera and action. The same person can appear in a room that was in no input.
Vidu’s API documentation states that the endpoint “accepts 1 to 7 images,” with image and text subjects capped at 7 in total. Kling calls the equivalent asset an Element — “a composite asset” built from multiple angles — and binds up to 3 elements to a start frame or start-and-end pair in the 3.0 model.
Which models offer it?
Vidu (Shengshu) exposes reference-to-video across its Q3, Q2, Q1 and 2.0 models; Q3 adds audio and 3–16-second clips at 540p to 1080p. Kling 3.0 supports up to 7 reference characters in video. Veo 3.1 calls the feature Ingredients to Video, uses 1–3 reference images and marks output with SynthID. Model entries: Vidu, Kling, Veo, Wan.
Why is identity locking a consent issue?
Image-to-video degrades a likeness over time; reference-to-video preserves it. That makes it valuable for film work and dangerous for intimate content: a few ordinary photos can place a real person into a scene they never took part in. Under the US TAKE IT DOWN Act such a clip is a “digital forgery” when a reasonable person cannot distinguish it from an authentic recording, and publishing it is a federal offense. Under the EU AI Act it is a deepfake that must be disclosed and machine-readably marked from August 2, 2026; from December 2026 systems built to produce non-consensual intimate content are prohibited in the EU.
How do we treat it in reviews?
No undress app in the current ranking exposes a multi-reference video mode; the video features tested in Ainudez, Deep Undress and Pornworks are single-image I2V, and Pornworks adds face swap on existing footage. When a reviewed tool adds reference-to-video, the editorial team tests it with synthetic or consenting-adult inputs only, records whether it checks consent for uploaded references, and notes output marking — items feeding the privacy score in how we rate. Victims of an identity-locked clip can hash it with StopNCII.org or report it here through the NCII form. This page is general information, not legal advice.
Sources
All sources accessed September 3, 2026.
- Vidu API — reference-to-video, 1–7 images
- Kling Element Library user guide — Kling 3.0
- Veo — Ingredients to Video, SynthID — Google DeepMind
- Wan2.2 — S2V-14B reference model — GitHub
- TAKE IT DOWN Act, S.146 — “digital forgery” — Congress.gov
- Regulation (EU) 2024/1689, Articles 3(60), 50 — EUR-Lex
Corrections: editorial@clothoff.ai · Editorial policy
Frequently asked questions
How many reference images does reference-to-video need?
Vidu accepts 1 to 7 images per request across its Q3, Q2, Q1 and 2.0 models. Kling 3.0 binds up to 3 elements to a start or end frame and supports up to 7 reference characters in video. Veo 3.1 uses 1 to 3 reference images. More angles generally give a steadier identity across the clip.
Is reference-to-video the same as face swap?
No. Face swap edits an existing video by replacing one face. Reference-to-video generates a completely new clip in which the referenced subject appears, including body, clothing and motion. The second is harder to detect because no original footage exists for comparison.
Do undress apps offer reference-to-video?
Not in the tools tested so far. Ainudez, Deep Undress and Pornworks provide single-image image-to-video, and Pornworks adds face swap on the Ultimate plan. Wavespeed lists general video models via API. The editorial team re-tests video features at each review update and records any change in the review.
Can a reference-to-video clip be identified as synthetic?
Veo output carries SynthID, and hosted services increasingly attach C2PA metadata. Open-weight pipelines leave no mandatory mark, so detection relies on forensic classifiers or platform hashing. From August 2, 2026 the EU AI Act requires machine-readable marking of synthetic video by providers serving EU users.
Is it legal to put a real person into a generated video?
With that person’s consent, yes. Without it, an intimate clip is a “digital forgery” under the TAKE IT DOWN Act, with up to 2 years in prison for adult subjects and 3 for minors, and covered platforms must remove it within 48 hours. Non-intimate uses can breach publicity or state deepfake laws.
Why does this site care about identity locking?
Because the privacy score in every review weighs how a tool handles other people’s likenesses. A mode built to preserve a specific person’s identity across scenes needs a consent check on uploaded references and output marking. Tools that skip both lose points, and reviews say so.