Text to speech
Type anything, pick a voice, hear it narrated โ no per-message length cap, just your plan's monthly character allowance.
Generation usually takes under a minute, but can take up to a few minutes after a quiet period while the voice engine wakes up.
Clone any voice
Record or upload ~10-20 seconds of a voice, then type what it should say.
Generation usually takes under a minute, but can take up to a few minutes after a quiet period while the voice engine wakes up.
Video
Four ways to make a video with Lucy - pick what fits.
- Your video, hyper-realistic. Your own photo/video + a script - exactly your face, powered by Kling.
- Cinematic. Your photo + a scene you describe - Veo generates the shot.
- Pick a character. 5 ready-made AI actors, always the same face.
- Pay as you go. Any prompt (+ optional photo/audio), any engine, no subscription.
The first three modes share one Video-plan balance: 40 credits/month. Whichever mode you use, your photo, video, and any audio are sent to third-party AI vendors (Kling, Veo, Seedance, and the fal.ai platform we use to reach them) for processing.
Your video, hyper-realistic
Upload your photo or a short video of yourself, type what to say - we animate exactly your face to say it.
Example: a real photo, dubbed with a Lucy voice via Kling.
Powered by Kling - the only engine in our tests that reliably keeps your exact face, not a lookalike.
Cinematic
Your photo + a scene you describe - Veo generates the shot around it.
Example: the moon-surface scene, from the prompt below.
A real limitation, not hidden: the more your reference photo moves within the scene, the more the face can drift from your real one - Veo regenerates the whole scene rather than animating your exact photo.
Pick a character
5 ready-made AI actors, always the same face - tap one to hear them, then type what they should say.
Example: Harper, one of the 5 characters below.
Comes from your Video plan's 40 credits/month - each video costs 8.
Pay as you go
Any prompt, plus an optional photo/video and audio - pick your engine, no subscription.
Sign in to buy video credits and generate.
Same flat price per video regardless of engine - real clip length differs (Kling is a hard 5s, the other four are 8s). Every engine gives you real lip-sync when you add audio: Kling does it in one step (needs a photo); every other engine renders the scene first, then a separate lip-sync pass matches the mouth movements afterward.
AI video models
We ran the same reference photo and the same prompt through five video engines - tap one to see the result.
Pay as you go above lets you generate with Kling, Veo, Grok, MiniMax, or Seedance. Quality, speed, and reliability vary a lot by engine and by scene - here's the identical lunar scene, run through each one, so you can see the difference before you pick an engine to generate your own.
Our pick: the best face consistency and scene quality of the five we tested. Note: this is a newer Kling version than the "Kling 2.1 Master" you can actually generate with above.
Caveat: these models change constantly - the labs and fal ship new versions often. This is a snapshot of where each one stood when we tested it (2026-09-12), not a permanent ranking.
That's the exact prompt we used above - want to try your own?
Just for fun: can AI promote your product?
We asked each model to put one of our AI characters in a real, spoken, two-scene ad for our own product - no distortion allowed.
Harper (one of our AI characters) actually speaks the script below, in her own voice, and changes both outfit and location partway through - a yoga mat to a corner office, holding the tumbler the whole time. Since no engine here can take two reference photos of a real person in one generation, we built each scene as its own separate video (its own reference photo, its own line of the script, its own lip-sync pass), then stitched the two scenes together afterward.
To be clear: this tumbler isn't a real Lucy Labs product - it's a relabeled stock photo, purely for testing how well each engine keeps a product's logo intact.


Try this with Harper + your cup
Use the real Harper face and Lucy Labs cup references above for a quick original product ad experiment.
- 1. Upload Harper face
- 2. Upload the cup
- 3. Paste the Old Spice style reference: youtube.com/watch?v=uLTIowBF0kE
- 4. Choose a model
- 5. Click Generate my video
That YouTube link is only a style reference. Lucy Labs creates an original ad from your brief and references, not a copy of the Old Spice commercial. When it finishes, download both the silent MP4 and the Kling-dubbed, lip-synced MP4.
Open the product ad builder

The two reference photos we actually used - one per scene, each combining Harper's photo with the tumbler photo.
Real lip-sync end to end (Kling's Avatar feature drives both scenes directly from the audio) - the strongest result of the three, crisp logo and a genuine outfit + location change.
The full script Harper says out loud:
โI'm on a yoga mat. Now I'm in a corner office, forty floors up. Anything is possible when your tumbler works as hard as you do. Check out Lucy Labs.โ
Those are the exact prompts we used - want to try your own version of each scene?
Scene 1 (yoga mat):
Scene 2 (corner office):
Want to make a multi-scene ad like this yourself?
Every engine above only accepts one photo per generation - none of them can jump between two scenes on their own. Here's exactly how we built the version above: (1) merge your two reference photos (e.g. yourself + your product) into one image using a free tool like fal.ai's Flux Kontext, Canva, or Photoroom; (2) generate one video per scene here, each with its own merged photo and its own prompt (different wording per scene, matching what's said in that part of the script); (3) stitch the resulting clips together with a free tool like Kapwing, Clideo, or CapCut.
Real-person policies explain the other two: Seedance blocks any photorealistic AI face outright, no matter the prompt - left out of this round entirely rather than re-spending credits to reconfirm an already-proven block. Veo flags prompts that read like a specific real person endorsing a named brand, so it can't generate Harper by name - but as the clip above shows, it will render a generic, unnamed person holding the real product once the named-identity part is dropped.
Caveat on the dubbing itself: watching all three side by side, Kling's lip-sync is clearly the best of the three - it generates the mouth movement and audio together in one pass. Grok and MiniMax's two-step process (silent video, then a separate lip-sync pass laid over it afterward) is real and does work, but the mouth-to-word match is noticeably less convincing than Kling's. If a spoken, dubbed ad is what you actually need, Kling is the one to pick today.
Honest caveat: a snapshot from 2026-09-12, not a permanent ranking - these models change constantly.
Can AI promote your product?
Make a cinematic ad from two images
Add your product, choose who appears, and tell us what the ad should feel like.
Selected: No product image yet
