Marketing

AI reels for a real business: brief in, video out

A reel for a small business is about nine seconds of footage, four lines of on-screen text and a piece of music. A video model can make the footage from a written brief in roughly the time it takes to boil a kettle. Here are four that Velofound made for a one-van bike repair shop in Denver — the brief behind each, how long the render took, and what had to be fixed afterwards. Then a fifth that didn't work, which is the one worth reading, because the reason it failed will be the reason yours fails too.

MarketingUpdated September 10, 2026By the Velofound team

"AI video" covers two quite different things. One is a model that invents a scene from a sentence — text to video. The other is a model that takes a still image and animates it — image to video. For a business, the second is almost always the right one, and the reason is worth understanding before you brief anything: a text-to-video model has never seen your business, so it invents a generic version of it. Seed it with a still that already looks like your street, your kit and your light, and the motion happens inside a frame that was yours to begin with.

So the chain runs: a written brief becomes a still image, the still becomes eight to twelve seconds of motion, on-screen text and a caption go on top, and the result lands on a day in the content calendar. Roughly three minutes of machine time and about ten minutes of yours, most of it spent deciding whether you like it.

The honest expectation, before any of the examples: this produces b-roll. Atmosphere, motion, a mood, a surface for text to sit on. It does not produce documentary footage of your business, it will not show a real customer, and it cannot demonstrate that you are good at your job. What it can do is stop the account being empty, which is the actual problem most small businesses have.

Four reels, and what each one cost in time

One build writes five creative briefs — three photos and two reels — so four reels is two builds, three weeks apart, which is why the brief numbers repeat. The two from the September build are the reels on days 3, 4, 11 and 12 of the fourteen-day calendar. Each block below is the full brief as it was written, the render times as they came back, and what happened to the file before it went up. The times are Northline's own runs on whichever video model the provider had available at the time; they are not a benchmark, and a busy queue roughly doubles them.

The 40-minute tune-up · September build · brief b2 · 9s · Instagram Reels, TikTokExample

What it's for. Make the forty-minute promise concrete for someone who has never had a bike fixed anywhere but a shop.

The prompt. Slow push-in on a road bike upside down on a portable repair stand in a suburban driveway, low golden evening light, a cargo van softly out of focus behind it, no hands in frame, calm and unhurried, photorealistic, shallow depth of field, no text, no logos.

On-screen text, in order.

  1. The 40-minute tune-up.
  2. In your driveway.
  3. Gears. Brakes. Chain. Tyres.
  4. $66. Saturdays go first.

Audio. Something unhurried and instrumental — picked from Instagram's own audio library at upload, not supplied with the file.

Render. Still 24s. Video 2m 12s on the first attempt, 2m 31s on the second. Second kept.

What was fixed afterwards. The first attempt had the van sliding a foot to the left with nothing driving it, so it was re-rendered with 'the van is parked and stationary' in the prompt. The keeper still lost its last 1.4 seconds to a trim, where the repair stand's legs drifted. On-screen text was added afterwards in the phone's own editor rather than baked into the render — text asked for inside a generated video comes out warped more often than not.

Three noises · September build · brief b5 · 12s · TikTok, Instagram ReelsExample

What it's for. Answer the question customers actually ask, with the information carried entirely by on-screen text.

The prompt. A rear bicycle wheel spinning slowly in shallow focus against a dark workshop background, single hard side light catching the rim, gentle continuous motion, photorealistic, no hands, no text.

On-screen text, in order.

  1. Three noises, three problems.
  2. Click under load → bottom bracket.
  3. Squeal at the lever → glazed pads.
  4. Grind while pedalling → worn chain.
  5. All three, same day.

Audio. Something with a beat, so the text beats can be cut to it.

Render. Still 20s. Video 1m 39s on the first attempt, 1m 52s on the second. Second kept.

What was fixed afterwards. The first attempt spun the wheel backwards, which nobody consciously notices and everybody finds slightly wrong. Re-rendered with 'wheel rotating forwards, top of the rim moving away from camera' added to the prompt. This is the useful category of failure: cheap to spot, cheap to fix, and the fix is a sentence.

Where we were this week · August build · brief b2 · 8s · Facebook, Instagram ReelsExample

What it's for. Show the service area as a place rather than a list of postcodes, under a caption that lists the actual week's jobs.

The prompt. Slow forward tracking shot down a tree-lined residential street on an early autumn morning, low sun through leaves, parked cars, wide pavements, empty of people, photorealistic, handheld feel, no text, no logos.

On-screen text, in order.

  1. Nine bikes this week.
  2. Wash Park. Sloan's Lake. Baker.
  3. We come to you.

Audio. Ambient, low. This one is watched with the sound off more than any of the others.

Render. Still 19s. Video 1m 48s. Kept first time.

What was fixed afterwards. Nothing in the footage. The caption did the work — the reel is atmosphere, and every specific claim (nine bikes, the three neighbourhoods, the tandem that hadn't moved since 2019) sits in the text, where it can be true.

Before the first frost · August build · brief b5 · 10s · Instagram ReelsExample

What it's for. Sell the $48 winter check in September, while there's still time to book it.

The prompt. Rain beading on the top tube and brake lever of a parked bicycle, very early morning, flat grey light, water running slowly off the frame, extremely shallow focus racking along the tube, photorealistic, no hands, no text.

On-screen text, in order.

  1. First frost is usually mid-October.
  2. Cold air drops your tyre pressure.
  3. Wet roads glaze your pads.
  4. Winter check: $48, on your street.

Audio. Quiet, slightly melancholy. Matches the light rather than the offer.

Render. Still 22s. Video 2m 31s. Kept first time.

What was fixed afterwards. Colour pulled down slightly — the model returned it warmer and sunnier than the brief asked for, which undercut the point of the post. Fifteen seconds in the phone's editor.

Add it up and the pattern is clearer than any individual reel. Four usable clips took six video renders and just under fourteen minutes of machine time. Two were kept first time; two needed one more attempt each, and both rejections were the same species of problem — something in the frame moving in a way physics wouldn't allow, a van sliding sideways and a wheel turning backwards. Every fix was a sentence added to the prompt or fifteen seconds in a phone editor. None of them required a camera.

The one that didn't work

The brief the founder most wanted was one they wrote themselves, over the top of a generated one, and it is the obvious idea: show the actual work. Hands laying a chain-checker gauge onto a chain, the gauge dropping in, the number readable. It is the shot that would have proved the thing every other post only asserts.

The chain measurement · September build · a sixth brief, written by the founder · abandoned after three attemptsExample

The prompt. Close-up of a mechanic's hands laying a chain wear indicator onto a bicycle chain, the tool dropping into the links, workshop light from the left, shallow focus, photorealistic, the number on the gauge legible.

Attempt one, 2m 05s. The hand had six fingers for four frames in the middle of the shot. Nothing else was wrong with it, which somehow made it worse — you watch it twice trying to work out what you saw.

Attempt two, 2m 18s, with "anatomically correct hands, five fingers" added. The fingers held. The chain didn't: the links changed size across the shot and the gauge merged into the chain about a second in, so the tool and the thing it was measuring became one object.

Attempt three, 2m 44s, pulled back to a wider frame to make it easier. The gauge came out with markings on it that were not numbers — the confident, wrong-shaped glyphs these models produce whenever text appears in a scene.

Seven minutes of render, nothing usable. The founder shot it on a phone the next morning, propped against a toolbox, in about forty seconds. That version is on the calendar.

This is worth generalising, because it is not bad luck and it will not be fixed by a better prompt. Video models are trained to produce something plausible frame to frame, and three things are systematically hard for that: hands doing precise work (many joints, self-occlusion, a shape everyone can check at a glance), repeating fine structure like a chain, a cassette, a keyboard, spokes, teeth — anything where the eye counts the elements and notices when the count drifts — and legible text inside the frame, which comes out as convincing letter-shaped marks that spell nothing.

Northline's brief asked for all three at once. That's not a hard shot, it's an impossible one, and forty seconds with a phone beat seven minutes of compute by a distance that isn't close.

The rule that comes out of it: generate the frame, film the craft. Anything that is atmosphere, place, light, weather, motion or a surface for text — generate it, and it'll be better than the stock photo you'd otherwise use. Anything where the point is that you did it — your hands, your workshop, the actual before and after, your face — film it, badly, on the phone in your pocket. Customers forgive shaky footage instantly. They do not forgive a hand with six fingers, and they cannot tell you why the video made them uneasy.

What actually happens when you press the button

  1. The brief is written first, and it is not a prompt

    A brief has a title, a goal, the described shot for the image model, four or five on-screen text beats, a spoken line if there is one, a duration, an audio suggestion and a three-item shot list. The described shot is the only part the model sees. Everything else is for you and for the caption, which is where the specifics live.

  2. A still is generated from the described shot

    An image model produces a single frame — around twenty seconds. This is the cheapest place to iterate, and the right one: if the still is wrong, no amount of motion rescues it. Regenerate here before you spend a video render on it.

  3. The still is animated

    The frame is handed to a video model along with the movement described in the brief. Ninety seconds to three minutes for eight to twelve seconds of footage, depending on which model is available and how busy it is. Image-to-video rather than text-to-video, for the reason at the top of this page: the shot stays inside a frame that already looked like your business.

  4. Text and audio go on afterwards, not inside

    On-screen beats are added in an editor over the finished clip, never asked for inside the render, because text generated inside a video comes out warped. Music is chosen in Instagram's or TikTok's own library at upload — trending audio can only be attached inside the app, and the brief's audio suggestion is a description of a mood, not a file.

  5. It lands on a day in the calendar

    The finished reel is attached to the post that uses it, with the caption, hashtags, platform and time already set. Then it waits for you — Velofound does not post to Instagram, TikTok or Facebook. You download the file, upload it, and tick the post off. That boundary is set out in full on the social media calendar page.

When video generation isn't switched on: you get an animatic instead of a blank box — the brief's on-screen beats timed over the generated still, which is a real, postable Story and a perfectly good way to test whether the words work before spending a render on the footage. It is not a reel, and it isn't described as one.

The specs, and three checks before it goes up

Get these right once and they stop being a decision. Everything below is current for Instagram and Facebook Reels as of September 2026.

Instagram and Facebook Reels technical specifications, September 2026
SettingWhat to useWhy
Aspect ratio9:16, strictlyAnything else gets pillarboxed and loses the full-screen frame that makes a reel a reel.
Resolution1080 × 1920Below this the platform's re-encode makes it visibly soft.
Length6–30 seconds for a businessInstagram raised the cap to three minutes in January 2025 and has taken longer since, but says Reels over three minutes aren't recommended to new audiences — and a service business rarely has thirty seconds of anything to say.
FileMP4 or MOV, H.264 video, AAC audio, 30fpsThe safe combination everywhere. 60fps only helps for genuinely fast motion.
Safe zonesKeep text out of the top ~15% and bottom ~20%That's where the username, the caption and the interaction buttons sit. Text under them is text nobody reads.
Cover frame420 × 654, picked by handThe default cover is whatever frame the platform grabs. On a grid it's the only thing anyone sees.
CaptionsOn, alwaysMost people watch with the sound off. A reel whose point is in the voiceover is a reel with no point.

Then three checks that take about a minute and catch most of what goes wrong:

  1. Watch it once with the sound off. If it stops making sense, the text beats are doing too little.
  2. Watch it once on a phone, at arm's length. Text you sized on a laptop is routinely too small, and the bottom line is routinely under the interface.
  3. Look at the hands, the wheels and any writing. Those three are where generated video fails, and you will spot it in a second once you know to look.

What to brief, and what to leave to your phone

Generate: place and weather
A street, a season, low light, rain on tarmac, a room with nobody in it. These are the shots you'd otherwise buy as stock, and a brief can put your own street's houses and light in the frame where a stock library can't.
Generate: motion under text
Anything whose job is to move gently while four lines of text carry the information. The 'three noises' reel works because the footage is a background, not a claim.
Generate: the product at rest
An object, well lit, still or slowly turning, no hands touching it. The failure modes all live in manipulation.
Film: hands doing the work
The measurement, the repair, the pour, the stitch. Forty seconds on a phone beats any render, and it's the footage that proves you can do the thing.
Film: your face, and your actual premises
A generated person who is meant to be you is uncanny at best and dishonest at worst. Nobody has ever unfollowed a small business for filming badly.
Film: before and after
The whole value is that it's the same object twice. Two generated frames are two different objects, and a customer who works that out has learned something about you that you didn't want to teach them.

The last thing, on disclosure, because it's a real question rather than a legal footnote. Images from the major models carry provenance metadata in the file, so Meta often attaches its own "AI info" label without being asked. TikTok expects you to switch on the AI-generated toggle yourself when the content shows realistic scenes or people. Do both, and say so in the caption when the picture is doing any persuasive work. The cost of saying "generated" is roughly zero. The cost of a customer working it out for themselves, on the post where you were asking them to trust you with their money, is not.

How this sits in Velofound: five creative briefs are written for your business in one build — three photos and two reels — and each one generates its image or its video from the brief, with the caption and hashtags already written and a day, platform and time on a two-week calendar. Reels are made for you; posting them is yours. Facebook and Instagram ads are the one thing published on your behalf, and those go out paused. The whole fourteen-day calendar →

Common questions

How long does one AI reel take to make?

About three minutes of machine time — roughly twenty seconds for the still, then ninety seconds to three minutes for eight to twelve seconds of footage — plus ten minutes of yours for the text, the music and deciding whether you like it. Northline's four usable reels took six video renders and just under fourteen minutes of compute in total. A busy queue roughly doubles the render times.

Is the footage real video of my business?

No. It's generated from a brief written about your business, so it looks like your kind of street, your kind of kit and your kind of light without being a recording of anything. That makes it good for atmosphere and useless as proof. Anything that has to be provably yours — the work itself, your premises, your face — film on a phone.

Why do some AI videos look wrong even when nothing is obviously broken?

Usually motion that physics wouldn't allow: a wheel rotating backwards, an object drifting a few centimetres with nothing pushing it, a shadow moving the wrong way. Nobody consciously registers these and everybody feels them. They're also the cheap ones to fix — describe the correct motion in the prompt and re-render, which is what happened to two of the four reels above.

What kind of shot should I never brief?

Hands doing precise work, repeating fine structures like chains, spokes, teeth or keyboards, and anything with legible text in the frame. Those three are where video models fail systematically, not occasionally. A brief that asks for all three at once — like the chain-measurement one on this page — is not a hard shot, it's an impossible one.

How long should a business reel be?

Six to thirty seconds. Instagram raised the Reel limit to three minutes in January 2025 and has accepted longer uploads since, but says Reels over three minutes aren't recommended to new audiences — and a service business rarely has more than fifteen seconds of anything worth saying. Nine seconds with four lines of on-screen text is a completely respectable reel.

Do I have to disclose that a video is AI-generated?

You should, and often the platform does it for you. Images and video from the major models carry provenance metadata, so Meta frequently adds its own 'AI info' label automatically; TikTok asks you to switch on the AI-generated toggle yourself for realistic content. Add a line to the caption too when the visual is doing persuasive work. It costs nothing and it removes the worst version of this, which is a customer noticing on their own.

Can I use a generated reel as a Facebook or Instagram ad?

Yes — Reels is one of the placements Meta ads run in, at the same 9:16 and 1080×1920. Bear in mind that an ad is watched more sceptically than a post, so anything slightly off in the footage gets more scrutiny, not less. And ads created by Velofound are published paused for you to approve before anything spends.

What happens if video generation isn't available on my plan or the model is down?

You get an animatic rather than an error: the brief's on-screen beats timed over the generated still. It's a real, postable Story and a decent way to find out whether the words land before you spend a render on the footage. It's not a reel, and it isn't presented as one.

Brief in, video out.

Describe your business once. Velofound writes five creative briefs — three photos and two reels — generates the image or the footage for each, writes the captions and hashtags, and puts the fortnight on a calendar. Free to start.

Start free →