Tyler Pratt← All articles
Content

How Faceless Creators Give Every Shot Enough Time Before You Edit — The Fix Is In The Narration, Not The Timeline

How long should each shot stay on screen in a faceless video?

2026-08-14 · 9 min read
Quick answer

A shot lasts as long as the line spoken over it. So your shot length is decided in the script, not the edit. Work out the spoken seconds of every line at your real words-per-minute, then compare it to the time that kind of visual needs to be understood — about 1.5 seconds for a plain photo, 4 or more for a labelled chart, 5 or more for a document or screenshot. Anything short is a shot nobody read. Fix it by adding a narration line that does real work over that shot, not by stretching the clip in the editor.

Key points

The length of a shot in a narrated video is not an editing decision. It is a script decision you already made without noticing.

Here is why. In a faceless video the voice runs continuously and the visuals change under it. So a shot stays up for exactly as long as its line takes to speak. Write a five-word line and that shot gets under two seconds, no matter how much work went into it.

That is fine for a photo. It is a disaster for a chart with five labels on it, because a viewer needs about five and a half seconds to read one of those.

A creator on r/NewTubers described the trap from the inside: "Often you can't show something until the narration has introduced it, but once you do... the script may have moved on. I found I often had to add filler dialog just so that the visuals get more time on-screen."

He was right about the fix and wrong about the material. You do need more words over that shot. They should not be filler.

Why does a shot last exactly as long as its line?

Because in a narrated video the audio is the spine and the visuals hang off it. You do not cut to a shot and then decide how long to hold it. You cut to a shot when the line about it starts, and you cut away when the next line starts.

So the durations were written the day you wrote the script. Every line is a stopwatch. A 12-word line at 155 words per minute is about 4.6 seconds. A five-word line is 1.9.

This is also why the problem is invisible while you write. In a document, line 3 and line 4 look the same size. On the timeline, one of them is worth two and a half times the other.

And it is why the editor is the wrong place to fix it. You can stretch the clip, but the voice has already moved on, so you get a silent hold. Silent holds read as a mistake to a viewer, not as a pause.

How long does each kind of visual actually need?

It depends entirely on how much reading the shot asks for. A photo needs recognition. A chart needs reading. Those are not the same task and they are not close in time.

These are workable minimums for a narrated video. They are the same numbers the tool uses:

  1. Plain photo or single object — 1.5s. Recognition only. Nothing to read.
  2. A person or a face — 2s. Slightly longer, because viewers read faces.
  3. Wide establishing scene — 3s. The eye has to travel before it settles.
  4. On-screen text card — 1s plus 0.25s per word. Four words is two seconds.
  5. Simple chart, two or three elements — 3.5s.
  6. Labelled chart, four or more labels — 4s plus 0.3s per label. A five-slice pie chart with percentages lands near 5.5s.
  7. Animated chart — 4s minimum. The animation has to finish or it never happened.
  8. Document, screenshot or newspaper — 5s plus 0.5s per highlighted element. This is the one creators underfeed the most.
  9. Table of data — 5s plus 0.4s per row.

Add about a second the first time a viewer sees a particular graphic. The second time you show that same chart it is already learned, and it can go by much faster.

What does it cost you when a shot flashes past?

You pay twice, and you never see either bill. The first cost is the shot itself. If your chart was on screen for 1.9 seconds, nobody learned the breakdown. The chart did no work. You paid for it in time or credits and it delivered nothing.

The second cost is worse. The viewer registers that something was there and that they missed it. That is a small friction, and it repeats every time it happens. A viewer in an r/NewTubers thread described how this reads from the outside: "the most noticeable thing are the vids where the narration doesn't match the video."

That is the sentence to sit with. Viewers do not describe it as a timing problem. They describe it as the video not matching itself, and they attach it to AI narration generally.

And nobody will ever leave you the comment that would let you diagnose it. No one writes "your chart was on screen for 1.9 seconds." They just stop watching, and you see a retention dip with no explanation on it.

Why not just hold the clip longer in the editor?

Because the voice does not stretch with it. That is the whole trap. Hold the chart for four extra seconds and you have four seconds of silence over a static image.

A pause works when the viewer is doing something during it. Silence over a chart the narrator has stopped explaining does not feel like a beat. It feels like the file broke.

There is a second reason, which is that the edit is the most expensive place to discover this. An editor described that job plainly: "b-roll heavy gaming content with voiceover sync is genuinely one of the more painful edits to sit through." Every fix at that stage costs you a re-render of the voice or a rebuilt sequence.

At the script stage the same fix costs you a sentence. This is the argument for doing the timing pass before you generate anything.

What makes a time-buying line different from filler?

Filler restates. A time-buying line does a job the visual cannot do alone. Same word count, opposite effect on retention.

There are only three jobs worth writing, and every good one is one of these:

  1. Direct the eye. Tell the viewer where to look inside the shot. "Look at the two on the left before I say a word about them." The pause is now the viewer working.
  2. Read it out. Say the numbers or labels the graphic is already showing, in the order you want them read. "Seven years and two months, down to four years and nine." That is not filler, it is the caption.
  3. Add one fact. A detail that only makes sense while this exact shot is up.

What disqualifies a line: "as you can see here", restating the sentence before it, or anything you could paste over any other shot in the video. If it would work anywhere, it is filler.

Does this change your video length?

Usually yes, and usually in the direction you wanted. Most scripts are short on shot time rather than long, so a full pass adds seconds rather than removing them.

On the sample script the tool runs on, four shots were short by a combined 13 seconds across ten lines. Scaled across a full 46-line script that was 7:12, the rewrite landed at 8:04.

That crosses the eight-minute mark, which is the threshold for mid-roll ads. It matters here mostly because of what most creators do instead: they pad the outro. A 50-second outro added to reach 8:00 is 50 seconds of your worst retention.

Buying the same 50 seconds back on shots that were already asking for them is the same runtime with the opposite effect. Your visuals get read, and the number goes up on its own.

How do you run the pass yourself?

You paste one prompt into Claude Code and it builds the tool. It comes pre-filled with a sample script and shot list, so it works on the first run and you can see the map before you feed it your own.

Then you paste your narration and your shot list in order, set your words-per-minute, and it does the arithmetic on every line, classifies every shot, and writes the fixes.

The read speed is an editable field at the top, because that is the one number that differs per channel. Change it and the whole map recalculates.

Grab it below — drop your email and the prompt is on the very next page.

Can you turn this into a side hustle?

Yes — and it is one of the simplest ways to make money with AI. You do not have to use this tool only for your own work. You can run it for other people and charge for it.

The play is simple. Local businesses already want Paste your narration and your shot list. Get the exact seconds every visual is on screen, every shot that flashes past before a viewer can read it, and the extra narration line that buys it the time — written for you, not filler. — they just dread producing it. You produce it in minutes, deliver a clean result, and bill for the outcome. Going rates run $500 a month per client.

The best part is the cost to start: a free prompt — it pays for itself on the first job. The tool does the heavy lifting in minutes, so your margin is high and you can take on more clients without more hours. To get your first client, reach out to a few local businesses you already know. Do one for free, show them the result, and ask who else needs it.

FAQ

What words-per-minute should I use?

Use your own. Take a finished video, count the words in that script, and divide by the runtime in minutes. Most narrated explainers land between 140 and 170. If you have not measured yet, start at 155 and adjust once you have one render to check it against.

Is this the same as a shot list tool?

No. A shot list tool decides what to show. This assumes you already know, and checks whether each of those shots gets enough screen time to be understood. Different job, and it runs after the shot list rather than instead of it.

What if I have more shots than lines?

Then two or more visuals are sharing one sentence, and the tool flags them as stacked. Each one gets a fraction of a line and neither lands. It will tell you which one to drop or write the extra line so both get their own window.

Can I re-run it on every script?

Yes, that is the point. Enter your API key once in the Settings panel and run it on every video you write. It is a reusable tool, not a one-time output.

Written alongside the Narration Timing Fixer · More AI tools & articles