lynxr get started

The two-column script: SAY and DO side by side

A two-column script puts what is said beside what is shown, on the same row, so every spoken line has a planned picture. The left column holds the words. The right column holds the framing, the action and any on-screen text. Reading across one row tells you exactly what the viewer hears and sees at that moment.

What does a two-column script look like?

It is a table with one row per moment of the video and two working columns, SAY and DO. SAY is every word that comes out of your mouth. DO is everything the camera sees while those words are said. The rows run top to bottom in the order the video plays, and each row is labelled with the part of the script it belongs to.

Those labels come from the four parts in a UGC script template that holds up: the hook, the beats, the turn and the close. That page explains why SAY and DO are written apart. This one is the page you fill in. Copy the blank version below into a spreadsheet, a notes app or a sheet of paper folded down the middle.

PartSAYDO
HookThe first line, word for word.What the viewer sees before that line finishes.
BeatThe next thing that happens, in a sentence you would say out loud.Framing, action, and any on-screen text.
BeatThe step after that. A new step, not a repeat.A picture that changes from the row above.
TurnThe line where the video becomes about the product.The product doing something, not just held up.
CloseOne line that ends it, with one ask at most.The last shot, and where the camera stops.

Why do the rows matter as much as the columns?

Each row is a unit of time. When you read a row aloud, the length of the SAY line is roughly how long the DO line has to stay on screen. A row with a long spoken line and a one-word DO line, such as “talking”, is a stretch where nothing on screen changes. A row with a short line and three actions in DO is a row you won't be able to film in the time the words take.

That makes the table a pacing check you can run before you pick up a phone. Read down the DO column on its own. If three rows in a row say “to camera”, the middle of the video is one static shot, and you can fix that on paper instead of in the edit.

What goes in the DO column?

The DO column holds three kinds of information: framing, action and on-screen text.

  • Framing is where the camera is and how close it sits: phone propped on the counter, waist up; over your shoulder at the screen; hands only.
  • Action is the one movement that happens during the line: peel the lid, tilt the tub to the lens, put the phone face down.
  • On-screen text is any word that appears on the video but isn't spoken, written out exactly as it will appear.

Write DO lines you could hand to someone else to film. “Show the product” is not a DO line, because it doesn't say how. “Hold the bottle next to your face, label forward, for the whole line” is one.

What goes in the SAY column?

Only the words you will speak, written the way you'll say them. No stage directions, no “(excited)”, no notes to the brand. If a line needs a tone, put the tone in DO as something the viewer can see: lean in, lower the phone, laugh before the line. A clean SAY column can be read straight through as a voice track, which is the fastest way to hear whether the video makes sense without its pictures.

When a brand requires a line word for word, put it in SAY exactly as supplied and mark the row so the brand can find it. Where those lines should sit is covered in how to turn a brand brief into a script.

What if a line has no picture?

A line with an empty DO cell is the most useful thing the layout shows you. It means one of two things: the line is doing work the video doesn't need, or it is doing work that should be shown instead of said.

Try the fixes in that order. First, delete the line and read the row above straight into the row below. If nothing is lost, it stays deleted. If something is lost, ask whether the viewer could see it instead. “It's really easy to set up” has no picture. Your hands clicking the part into place does, and once the picture carries the ease, the spoken line is free to say something the picture can't, like where you keep it.

A filled two-column script

Here is the same table filled in for one short video. It was written for this guide, for a set of reusable silicone food bags, and the brand is left unnamed.

PartSAYDO
Hook“I stopped buying plastic sandwich bags, and it had nothing to do with the planet.”Close on your face at the kitchen counter, holding an empty box of disposable bags upside down. Text: sunday lunch prep.
Beat“That box ran out every single week, on the one night I actually prepped.”Shake the empty box over the counter. Nothing falls out.
Beat“So I tried the washable ones, fully expecting to hate cleaning them.”Pull a silicone bag off the dish rack and open it towards the lens.
Turn“What sold me is that they stand up on their own while you fill them.”Set the bag down. It stays upright while you spoon rice in with one hand.
Close“If you prep on Sundays, the ones I use are linked.”Seal the bag, drop it in the fridge door, shut the fridge.

Read the DO column alone and you get five different pictures in a row. Read the SAY column alone and you get a story that makes sense with your eyes closed. When both of those are true, the script is ready to film.

Do it with lynxr

If you would rather start from a video that already works than from an empty table, lynxr fills in both columns for you. Paste a public TikTok or Instagram link on the homepage and lynxr works out the format — hook, beats, turn, close — and writes it as a script for the company you make content for. Free to start: 25 scripts, no card.

Keep reading