· 11 min read
Rocky & Hardev: Every Prompt Behind Two AI Pahadi Superstars
Two AI characters from Manali and four reels, made in one day for $30.29 of generation. Every prompt, reference image, setting, cost and mistake, open-sourced so you can make your own.
Rocky Rohtang drives a taxi over the Rohtang Pass, and his hair is taller than his face. Hero Hardev is Manali’s self-declared superstar. Neither of them exists, but both of them are on Instagram.
I made both of them, and four reels, in one day on FLORA with Seedance 2.5. The generation bill came to $30.29. This post is the whole production archive: every prompt, the four reference images, the model settings, what each render cost, and the mistakes that cost me money.
The idea comes from Higgsfield, which open-sourced every prompt and asset behind its feature film Hell Grind. Mine is a much smaller film. The rule is the same: if it worked, you should be able to copy it.
How I worked
I directed; Claude operated. Claude (in Cowork) was connected to FLORA’s MCP server, so it wrote the prompts, uploaded the references, launched renders on my FLORA canvas and pulled frames for me to review. I made the creative calls: who the characters are, which takes to keep, which scenes to remake, and what looked fake.
Every render below ran at 720p, 9:16, with the bitrate set to high.
Step 1: Cast the characters on video, not stills
My first character failed. He was a Dubai finance bro, built from image-model stills and animated with motion control, and every clip looked like AI. A face made as a still and animated later looks pasted on.
So I cast the next one the way a film does: a five-second screen test from Seedance 2.5 text-to-video. White backdrop, close-up to full length, one raised eyebrow. Then I pulled a face frame and a full-body frame from each clip with FLORA’s free extract-video-frames action. Those two frames per man are the whole identity system.
Here’s the twist. I ran one character description twice: once “on a cinema camera”, once “on an iPhone by a casting assistant”. I got two different men. I liked both, so one character became a duo.
Screen test. Seedance 2.5 text-to-video (t2v-gengateway-seedance-2-5-t2v), 5 s, no references. $2.457 per take.
A five-second casting screen test, filmed on a cinema camera against a seamless white studio backdrop. Soft, even key light from the front left with a large fill, true-to-life colour, no colour grade. The shot opens on a tight close-up of the man's face, then the camera slowly and smoothly dollies straight back until he is seen full length, from the top of his hair to his shoes, by the last second.
The man is an Indian man from the Himachal hills, about forty, tall and broad-shouldered with a gym-built chest. He has warm wheatish skin with wind-chapped redness across his cheeks, a strong square jaw, thick dark eyebrows, a thin, neatly trimmed moustache, and an intense, completely deadpan movie-hero stare straight into the lens. His hair is absurd: a giant glossy blow-dried puff swept up and back to twice normal height, with one thick lock falling across his forehead, and a small traditional Kullu cap with a bright woven geometric band balanced on the very top of the puff. He wears a black leather biker jacket, open, over a chunky hand-knit cream woollen sweater with a colourful Himachali geometric band across the chest, faded blue jeans and white chunky sneakers, with gold aviator sunglasses hooked on the sweater's neckline.
He stands still, breathes, blinks naturally, turns his head slightly to his left and back to the camera, then slowly raises one eyebrow. Real human skin with visible pores and natural facial asymmetry: it looks like real footage of a real person, not CGI. No text, no logos.
The iPhone take is the same prompt with a new first paragraph:
A five-second vertical casting video filmed on an iPhone by a casting assistant, against a plain white paper backdrop in a small studio. Light from one large softbox on the left plus overhead office lights, slightly uneven, with a faint shadow on the backdrop behind him. Straight-off-the-phone colour, slight handheld drift. The video opens on a tight close-up of the man's face, then the assistant slowly walks backwards until he is seen full length, from the top of his hair to his shoes, by the last second.
These are the four frames every later prompt points at. Download them below and use them to learn the method. For your own reels, cast your own duo.
| Tag used in the prompts | Who | File |
|---|---|---|
@rocky-face.png | Rocky, face | download |
@rocky-full-body.png | Rocky, full length | download |
@hardev-face.png | Hardev, face | download |
@hardev-full-body.png | Hardev, full length | download |
In FLORA, each @tag is the exact filename of an uploaded reference. The real names are long IDs, so I’ve swapped in readable ones below.
Step 2: One prompt skeleton for everything
Every duo prompt after the screen tests has the same seven parts. This is the closest thing I have to Higgsfield’s “technical foundation”:
- Task. What happens, or “Recreate the scene in @scene.mp4 shot for shot”.
- Soundtrack. Which file is the only soundtrack, or “only engine and wind, no music”.
- Cast. “ROCKY is the man in @rocky-face.png and @rocky-full-body.png”, his role, then his hair and face in words, every time.
- Identity lock. Wardrobe, then “Keep each man’s face and hair exactly as in his own images, and never mix their faces.”
- Shots. A timestamped list: who is in frame, who speaks, whose mouth stays closed.
- Location. A real, busy, specific place.
- Look and bans. “Real footage… Not CGI. No text, no subtitles.”
Parts 3 and 4 do the heavy lifting. Parts 5 and 6 came from mistakes, which you’ll see below.
Record 1: Mall Road rotation
Reference. The “rotation” format from the @leopold.delarue × @lord.farquaad.sol collab (1.4M views). I used its 11.5 s audio as the beat.
Model. Seedance 2.5 reference-to-video (r2v-seedance-2-5), 12 s. Inputs: the four images and the audio.
Cost. $5.861.
A real vertical iPhone video on Mall Road in Manali at dusk. Two men take turns strutting towards the camera in a deadpan "rotation", every step and move landing on the beat of Audio 1, which is the only soundtrack.
Man A is the man in Image 1 and Image 2: a towering, flared, rigid tower of dark hair that is wider at the top, a heavy lock falling across his forehead, a tiny Kullu cap with a colourful woven band balanced on top, bushy eyebrows that nearly meet, wind-burnt cheeks and a thick moustache. Man B is the man in Image 3 and Image 4: glossy black hair swept into two big winged sides with a sharp lock across his forehead, the same tiny Kullu cap on top, a neater moustache. Both wear black leather biker jackets over cream hand-knit sweaters with a colourful Himachali pattern band, faded jeans and white chunky sneakers. Both keep completely serious, deadpan faces the whole time.
The phone sits still at waist height with the 0.5x ultra-wide lens, looking up the busy pedestrian street: small shops with woollen shawls and hanging signboards, a momo stall, warm yellow shop lights, a steep pine-covered mountainside behind with a thin dusting of snow on the ridge, tourists in puffer jackets and woolly caps strolling, a few stopping to stare and film on their phones.
Choreography: Man A struts straight towards the lens with a slow swaggering hip-sway walk, snapping his fingers on the beat, fills the frame, then swerves out past the right edge. Man B walks in from the crowd behind with the same strut and a smooth shoulder shimmy, stops close to the lens, raises one eyebrow, and exits past the left edge. Man A comes back, does a slow spin and a sharp point at the camera. For the last beats both walk in side by side, stop together right in front of the lens, and stare into it without moving.
Look: real phone footage, natural dusk light mixed with warm shop lights, mild ultra-wide distortion near the edges, slight sensor noise, real skin texture, the hair shapes staying rigid as they move. No text, no captions, no subtitles, no dialogue.
What happened. Mall Road at dusk looks real, and the rotation works. But in places the two faces blended into each other. This prompt calls the references “Image 1” to “Image 4”. With four unlabelled faces in one clip, the model mixes them. Every later prompt tags each file by name and says “never mix their faces”.
Record 2: Deewaar, “Mere paas maa hai”
The failed version first. I rendered one shot per man with his half of the dialogue ($3.438 + $2.457), stitched the two, and laid the full audio back on top. The pictures were great, but the lip-sync was a mess. Rocky couldn’t finish his sentence, and Hardev’s voice landed after his lips moved.
The fix: one generation for the whole scene. Seedance 2.5 video-to-video in reference mode takes the original scene as a blocking guide, plus both men’s images and the full dialogue. A timestamped shot list then tells it who speaks when. I measured the speech windows from the audio’s loudness in 100 ms steps: the first line runs 0.7–4.4 s, and the answer starts at 8.6 s.
Model. Seedance 2.5 video-to-video (v2v-seedance-2-5), task reference, 12 s. Inputs: the scene clip, the four images and the dialogue audio.
Cost. $6.778. Render time: about 18 minutes.
Recreate the scene in @deewaar-scene.mp4 shot for shot, following its blocking, camera angles, cuts and timing, but with two new actors and a new location. The dialogue is @deewaar-dialogue.wav and it is the only soundtrack.
ROCKY is the man in @rocky-face.png and @rocky-full-body.png. He replaces the man in the suit who speaks first. He has a towering, flared, rigid tower of dark hair almost as tall as his face, a heavy lock falling across his forehead, a tiny Kullu cap with a colourful woven band balanced on top, bushy eyebrows that nearly meet, wind-burnt cheeks and a thick moustache.
HARDEV is the man in @hardev-face.png and @hardev-full-body.png. He replaces the younger man who answers. He has glossy black hair swept into two huge rigid winged sides with one sharp lock across his forehead, the same tiny Kullu cap on top, a neat moustache and a strong jaw.
Both wear black leather biker jackets over cream sweaters with a colourful Himachali pattern band. Keep each man's face and hair exactly as in his own images, and never mix their faces.
Shots and lip-sync, matched to the audio:
0.0 to 4.5 s: a two-shot with BOTH men in the frame. Hardev stands on the left with his back three-quarters to the camera; Rocky faces the camera on the right. Rocky speaks his Hindi line from 0.7 s to 4.4 s, his lips moving on every word until the line ends, arrogant and deliberate. Hardev's mouth stays closed.
4.5 to 5.2 s: over Rocky's shoulder onto Hardev, who stays silent.
5.2 to 7.0 s: close-up of Rocky, silent and smug, waiting for an answer.
7.0 to 10.0 s: close-up of Hardev. He stays silent with glistening eyes until 8.6 s, then says his Hindi line from 8.6 s to about 10.4 s, lips exactly on the words, quiet and proud.
10.0 to 10.8 s: close-up of Rocky reacting, his smugness draining away, mouth closed.
10.8 s to the end: close-up of Hardev holding his gaze, mouth closed.
New location replacing the original set: night under an old stone arch bridge over the rushing Beas river in Manali, big rounded river boulders, deodar pines on the far bank, one warm sodium streetlamp on the bridge, thin mist over the water.
Look: real 1970s Hindi film footage shot on 35 mm, film grain, practical lamp light, deep night shadows, real skin texture, rigid hair catching the lamp light. Not CGI. No text, no subtitles.
What happened. Hardev’s lips open on “Mere paas” at 8.6 s and on “maa hai” at 9.0 s, right on the words, and the two identities stayed apart. There were two misses. The opening two-shot is so wide that Rocky’s lips are too small to judge. And the stone bridge looks like a set. Those gave me two new rules: faces go big whenever someone speaks, and real busy places beat invented ones.
One more thing: keep the model’s audio. Seedance re-renders the reference dialogue slightly, about 0.1–0.3 s earlier, and the lips follow its version, not yours.
Record 3: Yeh Dosti on the Rohtang road
This is the Sholay song scene, moved to the Manali–Rohtang highway. My two attempts with the song were both blocked. One used the film clip as the reference (v2v); the other used only the song’s audio (r2v). Both came back “flagged for possible copyright concerns”, one after about 18 minutes of rendering. Neither was charged.
So the final render has no song in it. The prompt asks for Rocky belting out an old film song, with engine and wind as the only sound. Seedance invented its own singing. On Instagram, I mute the clip and add the real track from the music library, starting at Rocky’s first open mouth (about 0.5 s).
Model. Seedance 2.5 reference-to-video (r2v-seedance-2-5), 14 s. Inputs: the four images, no audio.
Cost. $6.841.
Two best friends ride an old black motorcycle with a round sidecar along a mountain highway in Himachal. The driver is belting out an old Hindi film song with huge joy, mouth wide open on every line, while his friend laughs along.
ROCKY is the man in @rocky-face.png and @rocky-full-body.png. He drives the motorcycle and is the one singing. He has a towering, flared, rigid tower of dark hair almost as tall as his face, a heavy lock falling across his forehead, a tiny Kullu cap with a colourful woven band balanced on top, bushy eyebrows that nearly meet, wind-burnt cheeks and a thick moustache.
HARDEV is the man in @hardev-face.png and @hardev-full-body.png. He rides in the sidecar. He has glossy black hair swept into two huge rigid winged sides almost as tall as his face, one sharp lock across his forehead, the same tiny Kullu cap on top, a neat moustache and a strong jaw.
Both wear black leather biker jackets over cream sweaters with a colourful Himachali pattern band. Keep each man's face and hair exactly as in his own images, never mix their faces, and keep the hair rigid in the wind.
Shots:
0.0 to 5.0 s: a medium two-shot from the front of the moving motorcycle, both faces large in frame. Rocky drives and sings with big open-mouthed joy. Hardev turns to look at him and laughs.
5.0 to 11.0 s: a close-up of Rocky driving and singing, head swaying, one hand lifting off the handlebar for a flourish.
11.0 to 14.0 s: a two-shot from the side. Hardev leans back in the sidecar with his hands behind his head, grinning; Rocky keeps singing.
Location: the Manali to Rohtang highway on a bright sunny day, a winding road with hairpin bends, patches of old snow on the verges, deodar pines, the Beas valley and snowy peaks behind, colourful prayer flags, a couple of tourist taxis passing.
Sound: only the motorcycle engine, wind and road ambience. No music.
Look: real footage filmed from a camera car driving alongside, natural daylight colours, light film grain, wind moving their jackets, real skin texture. Not CGI. No text, no subtitles.
What happened. This is the most real of the four. There’s snow on the verges, prayer flags, and a state bus going past. Every rule from the first two records is in this prompt: named references, big faces while singing, a real road, timestamped beats. For Instagram I upscaled it 2.5× with Topaz Starlight ($4.397).
The rules, condensed
- Cast on video. Use text-to-video screen tests, then pull face and full-body frames. Don’t animate stills.
- One generation per scene. Per-shot renders with re-laid audio break lip-sync.
- Tag every reference by filename, say who is who, and add “never mix their faces”.
- Faces big when someone speaks. Keep wide shots for silent beats.
- Keep the model’s audio on dialogue reels. The lips follow its version.
- Real, busy places look real. Invented sets look like sets.
- Restate the hair in every prompt. Signature silhouettes shrink in wide shots.
- Timestamp everything: who is in frame, who speaks, whose mouth stays closed.
- Famous film songs get blocked. Render without them and add the track in Instagram. Dialogue passed.
- No real faces. No celebrity lookalikes, ever. The characters are original, and the remakes are fan tributes.
What it cost
| Render | Model | Length | Cost |
|---|---|---|---|
| Two screen tests | Seedance 2.5 t2v | 2 × 5 s | $4.914 |
| Mall Road rotation | Seedance 2.5 r2v | 12 s | $5.861 |
| Deewaar, split per shot (discarded) | Seedance 2.5 r2v | 7 s + 5 s | $5.895 |
| Deewaar, one pass | Seedance 2.5 v2v | 12 s | $6.778 |
| Yeh Dosti | Seedance 2.5 r2v | 14 s | $6.841 |
| Two blocked song renders | Seedance 2.5 r2v, v2v | 14 s each | $0 |
| Total | $30.289 |
That works out to about $0.49 a second for text-to-video and reference-to-video at 720p, and about $0.56 a second for video-to-video. Renders took 6 to 30 minutes each. Upscaling is extra: Topaz Starlight at 2.5× cost $3.767 for the 12 s rotation and $4.397 for Yeh Dosti. My earlier attempts, including the dead Dubai character, cost about $6.70 before any of this.
FLORA gotchas that cost me time
- Browser-recorded clips can have a broken frame rate. A reference trimmed with the browser’s MediaRecorder came back as “30000.0 fps”, and Seedance refused it. Running FLORA’s reverse-video action twice re-encodes it to a normal clip.
- Actions need a video stream. An audio-only MP4 fails with “Failed to load video”. Convert it to WAV and upload that.
- stitch-videos needs at least two inputs, so it can’t re-encode a single clip.
- Copyright blocks arrive late. One came about 18 minutes into a render. It wasn’t charged, but the time is gone.
- extract-video-frames is free. Use it for every review. It’s how I checked lip-sync frame by frame.
Take it
The prompts above are copy-paste ready, and the four reference images are in the table. The prompts and the method are yours to copy. Rocky and Hardev are my characters, so please cast your own duo rather than reusing their faces.
What’s not in the pack is the film clips and film audio I used as references. They aren’t mine to share. For a remake, act out the blocking on your phone and use that as the reference video, or use clips you have the rights to.
Rocky and Hardev post a new scene every week on Instagram: @rocky.aur.hardev.