Georgii EmelianovGuides

SaaS Demo Video: How to Make One That Sells

Most guides assume you can already record clean footage and that the output is one file. Both assumptions are wrong, and both are why the video never ships.

Overhead photograph of a black smartphone and a paper storyboard card on an oak desk in soft daylight.

You shipped the feature three weeks ago. Sales keeps asking for something to send after the call, the pricing page still shows a static screenshot, and the ad account is running text. Everyone agrees you need a SaaS demo video. Nobody agrees on what it should contain, and the draft script has been open in a tab since Monday.

The reason it stalls is rarely the writing. It stalls because the two hardest questions get skipped: what footage you can realistically produce, and how many different videos you are actually on the hook for. Answer those first and the script writes itself in an afternoon.

The short version

  1. Pick one job the video has to do: convert on the site, advance a deal, or earn a click in an ad.
  2. Pick one use case that ends in a visible outcome.
  3. Choose your input — screenshots, a recording you can make cleanly, or footage you cannot capture yet.
  4. Script to the footage you will have, at roughly 150 to 200 words per 90 seconds.
  5. Produce one source cut long, then cut it three ways.
  6. Place each cut where its job lives, and measure drop-off.

What you'll need before you record anything

The list is shorter than the production agencies suggest, but the first item is non-negotiable and it is the one most teams skip.

A product state worth filming. Seeded data that reads as real: plausible names, realistic numbers, a populated dashboard. A demo of an empty account is a demo of nothing, and viewers register the fakeness of Test User 1 faster than they register your value proposition.

One flow that ends in an outcome. Not a feature. An outcome — the report exported, the invoice sent, the alert routed. The viewer needs to watch something finish.

A script target, not a script. Roughly 150 to 200 words of narration per 90 seconds of video. Write to that ceiling and the pacing takes care of itself.

Somewhere to put the file. A video host with retention analytics beats a raw MP4 on your CDN, because Step 5 depends on knowing where people stop watching.

You do not need a studio microphone, an agency, or a face on camera. Those are optimisations, and three of the four cuts you are about to make have no voiceover at all.

Step 1: Decide the one job the video has to do

A demo that sells does one job. The reason most demos convert badly is that they are asked to do three at once, and the compromise satisfies none of them.

The three jobs, and what each one demands:

Convert on the site. Autoplays, silent, in a loop, above the fold. Roughly 20 to 30 seconds. It has to be legible with no audio and no context, because the viewer arrived nine seconds ago and has not read your headline.

Advance a deal. Sent after a discovery call, watched deliberately, often forwarded to someone who was not on the call. Two to three minutes, voiced, and it can assume the viewer knows what your category is.

Earn a click. Paid social or a launch post. Fifteen seconds or less, captioned, and the first two seconds carry the entire result.

A video asked to convert cold traffic, brief a buying committee and win an ad auction will do none of the three. Pick the job first; the length, the audio and the framing all fall out of it.

Ask three questions in this order, and stop at the first yes:

  1. Is there a deal in the pipeline this month that stalls without a video? Build the sales cut first.
  2. Is paid traffic already running to a page with no motion on it? Build the site cut first.
  3. Neither? Build the site cut. It is the one that keeps working while you sleep.

Step 2: Choose your input before you write a word

Here is the question every guide on this topic skips: what can you actually film? Scripting before you answer it is how you end up with a storyboard that requires three test accounts, a staging environment, and a feature that is still behind a flag.

There are three honest inputs, and they are not equally good.

Screenshots you already have. The lowest-friction start. You have them from the App Store listing or the docs, and a tool can pan, scale and cut between them with motion that reads as intentional. The result can look genuinely good. It never shows your product moving, so it cannot demonstrate a flow — only a sequence of states.

A recording you can make cleanly. The default assumption of every article on the first page of results, and the right choice when it is true. It requires a build you can drive end to end without a crash, data that looks real, and the patience to redo the take when you fumble a click at 0:38. Budget three takes minimum.

Nothing presentable yet. The case nobody writes about. The UI is half-built, the flow needs credentials you cannot show, the animation only works on your machine. Most teams respond by waiting, and the video slips a quarter. The alternative is to have the footage generated: rebuild the flow as a working mock, drive it, and film that instead of your staging environment.

The input you have decides the video you can make. Choosing it last is why the script gets rewritten twice.

For a business-to-business SaaS with a gated product, the third case is more common than the first two combined. Nobody wants a live tenant's data in a public video, and scrubbing it by hand in post costs more than rebuilding the screen.

Step 3: Script to the footage, not the feature list

Write beats, not paragraphs. A beat is one thing the viewer watches happen, and the narration exists to name it — not to describe a capability the screen is not currently showing.

A four-beat spine covers almost every product:

  1. The situation, 15 to 20 words. The state the viewer recognises. No product yet.
  2. The action, 40 to 60 words. One flow, driven at real speed, ending in a visible result.
  3. The consequence, 25 to 35 words. What changed because that finished.
  4. The next step, 10 to 15 words. One instruction, one destination.

That is 90 to 130 words, which lands between 45 and 70 seconds of narration. If your draft is 400 words, the problem is not the writing — the scope is too wide, and the fix is to cut a use case rather than to talk faster.

Two rules that survive contact with production. Never narrate what is visible; if the screen shows a file uploading, do not say "now the file uploads." And never write a line that depends on footage you have not confirmed you can capture, which is why Step 2 comes first.

Step 4: Produce the footage

Three routes, matched to the three inputs. The differences that matter are mechanical, not aesthetic.

Animating stills. A still-based cut lives or dies on its motion design: a slow push toward the region that matters, a cut on the beat, and typography that carries the claim the screenshot cannot. The failure mode is drift, where the camera moves constantly and the viewer never gets a stable frame to read.

Recording and cleaning up. Record at the display's native resolution, then fix three things in post. Cursor jitter reads as nervousness, so smooth or hide the pointer. Dead air between a click and the response is where viewers leave, so cut it hard rather than speeding it up. And zoom toward each interaction, because a full desktop frame at 1080p renders your 13px labels illegible on a phone.

Generating the footage. Newer tools drive the product themselves and return finished video. That removes the take, the retake and the manual zoom pass in one step, and it is the only route that works when there is nothing presentable to record. It also constrains you to whatever runtime the tool can drive — which is the tradeoff worth understanding before you commit, and roughly what an AI app demo video generator actually does as a category.

Whichever route you take, produce one long source cut first. Do not produce three videos. Produce one, then cut it.

Step 5: Cut it three ways from one source

This is the step that decides whether the whole exercise costs you one week or three. The three jobs from Step 1 are not three productions. They are three cuts of the same source, and the edit list per cut is short and specific.

  1. Site hero. Trim to 20 to 30 seconds, strip the voiceover, remove the end card, and make the last frame match the first so the loop does not snap. Burn in two or three short captions, because it plays silent and unlabelled motion is decoration.
  2. Sales follow-up. Keep the full two to three minutes, keep the voiceover, add a title card naming the use case so a forwarded link makes sense to someone who missed the call. One CTA, at the end.
  3. Paid and social. Cut to the single strongest 15 seconds — usually beat two on its own. Reframe to 9:16, caption every line, and open on the moment of change rather than on your logo.
Three deliverables, one production. Teams that treat them as three shoots pay the cost three times and then update none of them.

Keep the source project, not just the exports. When beat two changes because the UI changed, you re-cut three files from one edit instead of rebuilding three timelines. If your tool represents the video as a spec file rather than a binary timeline, that re-cut is a diff:

# one recording, re-cut to two canvases without re-recording anything
reely effects source.mp4 --format vertical   --output hero.mp4
reely effects source.mp4 --format horizontal --output sales.mp4

Place each cut where its job lives: the hero above the fold, the sales cut in a link you can track per recipient, the short cut in the ad account. Then read the retention curve after two weeks. The cliff in the first cut tells you which beat to shorten in all three.


Where this goes wrong

Five failure modes account for most demos that get made and then quietly stop being used.

The tour. Nine features in 200 seconds, no job, no outcome. It happens when the script is written from the changelog. The tell is a narration full of "you can also."

The empty-state demo. Filmed on a fresh account because that was the fastest environment to reach. Every screen is a zero, and the viewer concludes nobody uses this.

The sales cut on the homepage. A three-minute voiced walkthrough autoplaying above the fold. It is the right video in the wrong slot, and its completion rate will look like a product problem when it is a placement problem.

Assuming sound. A large share of the views happen muted, in a feed or a shared tab. If the claim only exists in the voiceover, it does not exist.

The demo nobody can update. The one person who owned the project file has moved teams, so the video keeps showing last quarter's navigation. Producing from a spec rather than a timeline is the structural fix, and it is worth its own conversation.

Doing it without the manual work

If your product is a native iOS app, the generate route is unusually strong, and it is what Reely does. It is a Mac app plus an agent skill that runs locally — nothing uploads. You bring screenshots, a recording you already made on any stack, or nothing at all, and in the last case the agent rebuilds the feature as a working SwiftUI mock, drives it in the Simulator, and films it.

The mechanical part is worth naming, because it is what a screen recorder cannot do. When the agent drives the run, reely run writes a timeline.json sidecar of labeled interaction marks beside the recording, and the composition references those labels rather than timestamps. The cut, the zoom and the caption land on the tap instead of 200ms after it. Upload your own footage and there is no sidecar, so you are timing by eye like everyone else.

Now the honest fork, because most SaaS is not an iPhone app. If your product is a web app, the generate route above does not apply to you. Reely's device frames, canvas formats and brand extraction are all phone-shaped, and it cannot drive a browser. You still get the screenshot and existing-recording paths, but for having web footage generated, Slideshot is the better answer: its site states the agent drives your web app, priced at $0.90 per recording request with no subscription, checked 14 August 2026. That is how the two compare on web versus native, and for a browser-based product the recommendation is not Reely.

Frequently asked questions

How long should a SaaS demo video be?

Placement decides it, across a range of 15 seconds to 3 minutes. A silent site hero runs 20 to 30 seconds. A sales follow-up runs 2 to 3 minutes because the viewer opted in. A paid social cut runs 15 seconds or less. Producing one 90-second file and using it in all three slots underperforms in at least two of them.

What is the difference between a demo video and an explainer video?

An explainer video sells the idea; a demo video sells the product. An explainer uses animation or metaphor to establish why a problem matters and usually never shows real software. A demo shows the actual interface completing a real task. If your category is unfamiliar you need both, and the explainer comes first in the funnel.

How much does a SaaS demo video cost?

Agency production is typically quoted per finished minute and runs into the thousands, which is why most early-stage teams do it in-house. Doing it yourself costs a screen recorder, an editor, and roughly two days of your own time for a first version. Generated-footage tools sit between the two, priced per video or per month.

Do you need a voiceover or a person on camera?

No, and for two of the three cuts a voiceover is actively wrong. The site hero and the paid cut play muted, so their message has to be carried by on-screen captions. Record a voiceover for the sales follow-up, where the viewer chose to watch. A face on camera is optional and mostly matters for founder-led outbound.

Next

Do not start with the three-minute version. Cut the 20-second site hero first: one use case, one visible outcome, no voiceover, two captions. It is the smallest artifact that is useful on its own, it exposes whether your seeded data holds up on screen, and it is the source you will trim the other two cuts from next week.

One case is worth trying a generated run on: a native iOS app with nothing clean to record yet. If that is yours, download Reely for Mac and point it at the flow you would have filmed.

← All posts