← gavinpurcell.com
Prompt Recipe FLUX 3

How to make a 20‑second clip of a TV show that never existed

A reusable structure for period found-footage video prompts, plus one worked example that actually came out right. Built and tested on Black Forest Labs' FLUX 3.

The short version

FLUX 3 generates video with synced audio and dialogue, which means prompting it is closer to briefing a director than writing an image prompt. Keyword soup underperforms badly. Full sentences win.

Two things matter more than everything else combined, and both are counterintuitive:

Rule One

Open in the middle of something already going wrong. No greeting, no setup, no establishing shot. The first line of dialogue should be a reaction, not an introduction.

Rule Two

Name the defects, not the look. "VHS aesthetic" gets ignored. "A tracking distortion bar drifting up the frame" renders.

My first attempt at this opened on a woman standing still, about to say a line. It looked perfect and it was boring. Same set, same grain, same everything, but nothing was happening yet. Starting mid-action fixed it completely.

The template

Fill in the bracketed parts. Keep the structure. Everything is one continuous block of prose: no line breaks needed, no parameters panel, no settings to configure. Duration and aspect ratio ride along inside the text.

Template
A [YEAR] [FORMAT: local news broadcast / infomercial / sitcom / nature documentary], shot on [CAMERA: studio tube cameras / 16mm / a consumer camcorder] and dubbed to VHS: [THREE NAMED DEFECTS: soft interlaced footage, a tracking bar drifting up frame, red chroma bleed], 4:3. [LIGHTING: flat hard studio lighting, slightly overexposed]. The set is [SET, LISTING CHEAP MATERIALS BY NAME: foam board, laminate, a plastic ficus]. [CHARACTER: age, wardrobe, and above all HOW THEY PLAY IT]. The clip begins mid-event, already in progress, with no introduction.

0.0 to 7.0s: [CAMERA MOVE]. [ACTION ALREADY UNDERWAY: "he is already mid-swing", "she is already halfway through the slap"]. "[REACTION LINE, NOT A GREETING]"

HARD CUT.

7.0 to 13.5s: [NEW ANGLE]. [THE PHYSICAL GAG]. "[LINE]"

HARD CUT.

13.5 to 20.0s: [WIDE THAT REVEALS SOMETHING: the cheap set, a boom mic, an extra not listening]. [THE TURN]. "[BUTTON]" Hold.

Audio: [FIVE OR SIX NAMED LAYERS: room tone, 60-cycle hum, a specific mechanical sound, music character, where the voice sits in the mix]. No on-screen text, no captions, no modern camera moves, no cinematic color grade.

Why it's built that way

ElementWhat it's doing
Format line first Tells the model what this footage is before it worries about what's in it. Year plus medium does most of the work.
Three named defects Vague style words get absorbed and ignored. Specific physical artifacts render. Three is enough; five starts fighting itself.
Performance direction The highest-leverage sentence in the whole prompt. "Big smile, flat eyes, tiny pauses before the mean parts" beats three sentences of costume.
Timestamped beats Anything over ~7 seconds needs explicit time ranges with HARD CUT between them. Budget 5–7 seconds per beat, so 20 seconds is three.
Word count per beat Roughly 2.5 words per second including pauses. A 7-second beat holds about 15–18 words. Overwriting dialogue is the most common failure: the model speeds up delivery and the performance flattens.
Audio spelled out "Atmospheric ambience" produces nothing. "Boxy studio room tone, a faint 60-cycle hum, buzzy on-camera mic with hiss" produces the scene.
Negatives last Mostly to suppress on-screen text, which is the flakiest thing these models render, and modern cinematography, which they default to.

A worked example

Here's one that came out right on the first try. The whole clip is a guy protecting a lobster.

Prompt
A 1994 network sitcom, shot on video before a live studio audience and dubbed to VHS: soft interlaced footage, light tape hiss, mild chroma bleed, 4:3. Flat bright multi-camera lighting, that overlit sitcom apartment look. The set is a Chicago apartment with a green couch, a mustard refrigerator, and a bicycle on the wall. The clip begins mid-event, already in progress, with no introduction.

0.0 to 7.0s: Locked-off wide, proscenium framing. A gangly man in a flannel shirt is already falling backward over the arm of the couch, holding a live lobster above his head to protect it. He lands hard. The lobster is fine. From the floor, without getting up: "Gerald is fine. Everyone relax."

HARD CUT.

7.0 to 13.5s: Locked-off medium on a woman in a cardigan standing in the doorway holding groceries, staring at the floor where he is. "Who is Gerald."

HARD CUT.

13.5 to 20.0s: Locked-off wide. The apartment door opens and an older man in a bathrobe leans in without explanation. "Has anyone seen my lobster." He leaves. Neither of them moves. Hold on the closed door.

Audio: a real studio audience laughing and applauding at each beat, some laughs running longer than the joke deserves, boxy stage room tone, light tape hiss, faint 60-cycle hum. No on-screen text, no captions, no cinematic color grade.

What's doing the work here

Practical notes

Written up from a weekend of trial and error. FLUX 3 is Black Forest Labs' video model. Everything above was tested in their playground.