Brand & Content

AI-Generated B-Roll Graphics

How to produce broadcast-quality animated graphics for any video, using AI agents and code, without a motion designer and without learning to animate.

Status: Canonical · Living document Audience: Anyone who wants animated graphics for video. Content creators, agency owners, their clients, hobbyists, operators, and AI content agents. Technical skill is not required to run the workflow. What this is: The complete blueprint for a programmatic graphics pipeline. It explains what these animations are, what they are for, how they work, and exactly how to create them, from a single one-off graphic to a full studio producing many videos. How comprehensive: Detailed enough that an AI agent handed this document, a video blueprint, and a set of design stills can build the entire project repository from scratch and run the pipeline end to end.


How to use this guide

Read the Opening and Part I no matter who you are. They take ten minutes and give you the mental model that makes everything else obvious.

After that, you take one of two paths, which you choose in Part II:

  • One-off path. You need a single animation, or you already know exactly what you want. Skip the planning chapters and the studio architecture. Go Opening → Part I → Quickstart → Part IV (Design) → the minimal build in Part V.
  • Studio path. You produce videos repeatedly and want a system that compounds. Read it start to finish.

Every chapter is tagged [Both], [Studio], or [One-off] so you always know whether it applies to you.

The reference assets you will copy and paste (templates, commands, folder structure, schemas) all live in Part VII.


Table of contents

Opening

Part I. Understand it

  1. What this is, and what it is not
  2. Why it works and why it matters
  3. Who it is for, and everything you can make
  4. The mental model: four actors and one rule

Part II. Decide your path 5. Two paths: one-off vs studio 6. Before you begin (prerequisites) 7. Quickstart: one animation, fast

Part III. Plan it 8. The episode blueprint 9. Conversion-leverage tiering

Part IV. Design it 10. The brand system (design tokens) 11. Designing the stills with Claude Design

Part V. Build it 12. Setting up the studio (the repo) 13. The animation libraries 14. Building the components with AI agents 15. Preview, render, and handoff

Part VI. Operate it 16. The binding constraint and the anti-over-engineering rules 17. The compounding asset 18. When to productize and integrate

Part VII. Reference

  • A. The AGENTS.md template
  • B. The component build prompt template
  • C. Component family taxonomy and starter catalog
  • D. The shot list schema
  • E. Commands cheat sheet
  • F. Quality checklist
  • G. Troubleshooting
  • H. Glossary

Opening

Every video that looks like it has a team behind it has them. The counter that ticks up to a number and makes you lean in. The diagram that draws itself while someone explains it. The clean branded motion that tells you, before a single word, that the person on screen is serious.

For thirty years those graphics had a gatekeeper. A motion designer in After Effects, billing by the hour, taking days per video. If you could not afford one, your content looked like everyone else's. If you could, you waited, and you waited again for every revision.

That gate is gone.

You are about to learn how to produce broadcast-quality animated graphics for any video you make, on brand, in a fraction of the time and cost, without opening After Effects, without hiring a designer, and without learning to animate a single keyframe. And not with an AI video generator that invents a slightly different result every time you press the button. With something far more powerful, because it is far more controllable: graphics built from code, designed by you, and assembled by AI agents that do all of the technical work while you direct.

The numbers are not subtle. A thirty-second branded segment renders in under five minutes. A motion-graphics job that used to eat fifteen to twenty hours becomes an afternoon. And here is the part that compounds: every graphic you build makes the next one cheaper, because you are not making one-off videos, you are building a library that gets faster and better the more you use it.

Here is the shift most people miss. The bottleneck in great content was never the idea. It was the production. The reason a competitor's videos feel more credible than yours usually has nothing to do with their insight and everything to do with their polish, and polish has always been expensive. This guide hands you polish at near-zero marginal cost. When production stops being the constraint, the only thing left that decides whether your content wins is the thing you already control: having something worth saying, and proving it.

If you make videos for yourself, for clients, for your clients' clients, or for an audience, and you have ever looked at a slick animated explainer and thought "I could never make that," this is the document that proves you can. If you are building a studio, read every page. If you just need one animation today, jump to the Quickstart. Either way, by the end you will understand a capability that, a few years ago, did not exist.

Let's get into it.


Part I. Understand it

1. What this is, and what it is not [Both]

This guide describes programmatic graphic animation: producing animated video graphics by writing code instead of dragging clips and keyframes in an editor.

The engine underneath is Remotion, a framework that turns code into video. In a normal app, the screen updates when a user clicks something. In Remotion, the screen updates in response to a frame number. A thirty-second video at thirty frames per second is nine hundred frames, numbered 0 to 899. Your graphic is written as a recipe that says, in effect, "at frame 0 the number reads zero, by frame 90 it reads fifty-two." Remotion plays that recipe forward, takes a picture of the screen at every frame, and stitches the pictures into a video file.

In plain terms: a graphic is a description of what the screen should look like at any given moment in time, and the engine turns that description into a video. You never touch a timeline of clips. You describe, and it renders.

It is worth being precise about what this is not, because two confusions sink most people:

It is not an AI video generator. Tools that generate video from a text prompt are non-deterministic. Ask twice, get two different videos. You cannot control exact timing, exact wording, or exact brand colors, and you cannot reproduce a result. This system is the opposite. It is deterministic: the same inputs produce the same video, down to the pixel, every single time. When you need a diagram with the right labels in the right colors moving in the right order, you need precision, and precision is what this gives you. (AI does play a central role here, but in the decision layer, not the pixel layer. More on that in Chapter 4.)

It is not a replacement for a video editor or for human taste. It produces the graphics, the animated segments, the branded motion. A human still decides what to say, what to show, and how the final video is assembled, paced, and scored. Think of it as a precision parts factory, not a finished-furniture store. The factory cuts perfect parts on demand; a person still designs the piece and does the final assembly.

It is also not for everything. Content whose value is its rawness, a phone-shot vlog, a candid behind-the-scenes clip, a face-to-camera confession, should stay raw. Polishing it would undercut the very authenticity that makes it work. This system is for content whose value comes from clarity, credibility, and visual explanation. Knowing which is which is a creative judgment, and it is yours to make.

2. Why it works and why it matters [Both]

There are two reasons to care, one operational and one persuasive. Both matter.

The operational reason: it removes the production bottleneck and replaces it with an asset that compounds. Manual motion graphics do not scale. Every video starts from a blank canvas and costs the same fifteen to twenty hours as the last one. A programmatic library inverts that curve. The first video you make might require building five graphic types. By your tenth, most of what you need already exists and you are configuring rather than creating. By your thirtieth, you have a near-complete kit and the marginal cost of a new video's graphics has collapsed. You stop paying the full price every time. This is the difference between doing work and building leverage, and it is the whole reason to adopt a code-based approach over hiring it out clip by clip.

The persuasive reason: production quality is a credibility signal, and credibility is what converts. This is not vanity. Trust is the belief that someone will deliver in the future based on how they have behaved in the past, and on screen, your only evidence of how you behave is how your content looks and what it claims. When a creator says "we built a system that produces these results" and the screen shows a professionally animated diagram building itself with real numbers flowing through it, the visual is not decoration. It is proof. The polish is the argument. A claim spoken over a clean, precise, on-brand graphic lands as a demonstration; the same claim over a sloppy slide lands as a boast.

The two reasons reinforce each other. Cheap, fast, repeatable production lets you ship more proof, more often, on brand, which builds the trust that turns attention into customers. That is why a graphics pipeline is not a "nice to have" for a serious content operation. It is infrastructure.

3. Who it is for, and everything you can make [Both]

Who runs this. The workflow is built so that the person directing it does not need to write code. That makes it usable by:

  • A solo creator or founder producing their own content.
  • An agency producing content for clients (and teaching clients to produce their own).
  • A client's in-house team producing their own marketing.
  • A hobbyist who wants one good animation for a personal project.
  • An AI content agent, operating inside a larger system, generating graphics from an arbitrary video idea with no human in the loop at all.

The division of labor is the same in every case and it is the key to the whole thing: a human (or a directing agent) provides creative direction and approval; an AI coding agent does the technical building. You decide what the graphic should communicate and judge whether it looks right. The agent writes the code, wires it up, and renders it. You never have to know how the code works to get a result you are happy with. Hold onto that, because it is what makes everything below approachable.

What you can make. Because the engine can render anything a web browser can draw, the range is wide. Concretely:

  • Continuous animated explainers. A full video that is itself the animation, narrated over the top. Educational content, concept breakdowns, "how it works" pieces, science and finance explainers in the style of the polished explainer channels.
  • B-roll for talking-head video. The animated segments that cut in over or between a person speaking: counters, diagrams, comparisons, callouts. This is the workhorse use case for sales videos (VSLs), YouTube videos, course lessons, webinars, and keynote recordings.
  • Ad creative. Short, punchy, branded animated hooks for paid social and display ads, where the first second decides everything and brand-exact motion separates you from the feed.
  • Vertical social clips. Reels, Shorts, and TikToks in 9:16, often the same graphics re-rendered in a vertical frame (covered under multi-format below).
  • Podcast and audio visuals. Audiograms, animated waveforms, quote cards, and chapter graphics that turn audio into shareable video.
  • Product and SaaS demos. Animated feature walkthroughs, UI motion, onboarding explainers, and launch videos.
  • Data storytelling. Charts, dashboards, and metrics that animate, turning a report or a results page into a video that holds attention.
  • Pitch and investor videos. Animated traction charts, market diagrams, and roadmap timelines for fundraising and sales decks brought to life.
  • Industry-specific reels. Real estate listing animations, e-commerce product motion, event and webinar promos, testimonial and case-study videos, recruiting and culture videos, and more. If a video would be clearer or more credible with a clean animated graphic, this is the tool.

Any format. Graphics can be built to render in widescreen (16:9), vertical (9:16), or square (1:1), and a well-built component renders in more than one with little extra work, so a single graphic can be repurposed across YouTube, Reels, and feed posts.

The inputs are always the same two things. Whatever you are making, the pipeline takes two human-provided inputs and the agent handles the rest:

  HUMAN PROVIDES                         AGENT EXECUTES                        YOU GET
  ──────────────                         ──────────────                        ───────
  1. A blueprint outline    ─────►  builds the components,        ─────►  finished animated
     (what to show, in order,        wires them into scenes,              graphics: either a
     with rough timing)              renders the video                    continuous video or
                                          ▲                                discrete segments to
  2. Design stills            ──────────┘                                  drop into an edit
     (what each graphic
      looks like, as a picture)

Everything in Parts III through V is just an expansion of that diagram.

4. The mental model: four actors and one rule [Both]

Internalize four roles and one rule and the rest of this guide reads like common sense.

Actor 1, the engine (the clock). Remotion is the rendering engine and, more importantly, it owns time. It is the metronome. Every moving thing in your graphic gets its sense of "what time is it" from Remotion's current frame number.

Actor 2, the libraries (the tools inside). A handful of specialized libraries help build the visuals: one for drawing data and charts, one for choreographing complex multi-element motion, one for playing pre-made micro-animations, one for 3D. These are not separate video makers and they are not interchangeable with the engine. They are tools you reach for inside a single graphic to do a specific job. Most of the time you will not need any of them; the engine alone is enough. (Full detail in Chapter 13.)

Actor 3, the subject (the component). The subject is the meaningful thing on screen: the counter, the bars, the diagram, the org chart. In code this is a "component," a self-contained, reusable piece. The subject carries the data and the meaning, and it carries any effect that is physically attached to it, such as its own drop shadow or the glow coming off a number. If it moves with the thing, it belongs to the thing.

Actor 4, the environment (the stage). Everything around the subject that stays constant from graphic to graphic, the background, the brand color behind everything, a subtle vignette, a logo watermark, the safe margins, is the environment. It lives in a shared wrapper called the stage, and you build it once and reuse it for every graphic. The subject sits inside the stage. The subject does not know or care what is behind it.

This subject-versus-environment split is the single most useful structural idea in the whole system, so here is the test you will use constantly: if an element moves or changes together with the subject, it lives inside the subject; if it is constant background context that many graphics share, it lives in the stage. A card's shadow travels with the card, so it belongs to the card. The dark gradient behind everything is shared, so it belongs to the stage. The empty space around the subject is not something you draw at all; it is the result of layout and brand settings, governed by the stage and the brand system.

The one rule: the engine owns the clock. This is the law that keeps everything reliable. Any library or effect that moves over time must take its timing from Remotion's frame number. You never let another tool run on its own internal clock, because the moment you do, the output stops being reproducible, and reproducibility is the entire reason this approach beats AI video generation. Charts compute their shapes as still math and you reveal them by frame. Choreography libraries are built as paused timelines that you scrub to the current frame. Pre-made animations are seeked to the current frame. Everything bows to the metronome. Hold this rule and you will never produce a graphic you cannot rebuild.


Part II. Decide your path

5. Two paths: one-off vs studio [Both]

Two kinds of people use this guide, and they should not take the same route.

The one-off path is for you if you need a single animation, you are trying the system for the first time, or you already know exactly what you want. You do not need a blueprint, a tiering exercise, or a multi-project repository. You need the shortest line from idea to rendered graphic. Your route: read Part I, do the Quickstart (Chapter 7), design your still (Chapter 11), and run the minimal build at the end of Chapter 14. You can be done today.

The studio path is for you if you produce videos repeatedly and want the compounding asset described in Chapter 2: a growing library, faster every month, that eventually a team or an AI agent can operate. You will set up a proper repository structure, batch your work, and govern the system so it stays healthy as it grows. Read everything.

A useful way to decide: build the studio when you expect a third project. One graphic is a one-off. The second can be a one-off too. By the third, the structure pays for itself, and not before. Do not build a studio for a single video; that is over-engineering, and this guide will warn you away from it more than once.

6. Before you begin (prerequisites) [Both]

You need a small set of tools. They install once. The AI agent handles the technical parts, so do not be put off by the names.

  • A computer, Mac or Windows.
  • Node.js (version 18 or higher). This is the runtime that lets the video engine run on your machine. One-time install from the official Node.js site. If a command later says "node is not recognized," this is the missing piece.
  • An AI coding agent: Claude Code. This is the tool that writes and runs the code for you, from the command line. It requires an Anthropic plan. This is your builder; you will be giving it instructions in plain English.
  • Claude Design, for creating the visual stills (Chapter 11). This is where you decide what each graphic looks like before any animation is built.
  • A code editor (optional): VS Code. Helpful for looking at files, but the agent does its work in the terminal, so you can skip this if you prefer.
  • Git (optional, recommended for the studio path). This is version control. The agent can manage it for you; on the studio path it also enables a parallel-build trick covered later.

That is the entire toolkit. Notice what is not on the list: After Effects, a design degree, or the ability to write code.

7. Quickstart: one animation, fast [One-off]

This is the shortest path to a single rendered graphic. It assumes the prerequisites above are installed.

  1. Open a terminal and create a project. Type:
    npx create-video@latest my-graphic
    Choose the blank template, choose TypeScript, and let it install. Then enter the folder:
    cd my-graphic
  2. Teach the agent the engine. This loads the official engine instructions into your AI agent so it writes correct code:
    npx skills add remotion-dev/skills
  3. Design your still. In Claude Design, create a single picture of what you want the finished graphic to look like at its most complete moment. (Chapter 11 explains how to make a good one.) Save the image.
  4. Start the agent and describe the graphic. Run:
    claude
    Then give it the still and a plain-English description, for example:

    "Build a Remotion graphic that matches this still. It is a number that counts up from 0 to 52 and then shows '%' with the label 'positive reply rate' underneath. Total length 4 seconds at 30 frames per second, 1920 by 1080. The number should count up smoothly and settle, not bounce hard. Start simple, show me a first version, then we will refine."

  5. Preview it. The agent will tell you to run:
    npm run dev
    This opens a preview in your browser where you can scrub the timeline and watch the animation.
  6. Refine in the same conversation. Keep talking to the agent: "slow the count so it lands at 3 seconds," "make the label fade in after the number settles," "the brand blue is too bright, use this hex." Quality compounds across turns; do not try to get everything in one prompt.
  7. Render the final file. When you are happy:
    npx remotion render <composition-name> out/my-graphic.mp4
    The agent will give you the exact composition name. Your finished graphic is now in the out folder.

That is the entire loop for a single graphic. The studio path that follows is the same loop, organized so it scales and compounds.


Part III. Plan it

8. The episode blueprint [Studio] (one-offs may skip)

Before you build graphics for a video, you need to know what graphics the video needs. That document is the episode blueprint: the single source of truth for one video. At minimum it answers, for the whole piece, who it is for, what it is trying to make the viewer believe or do, and then, scene by scene, what appears on screen, in what order, for roughly how long.

For the build pipeline, the blueprint produces one critical artifact: a shot list, an ordered inventory of every graphic the video needs, each with a name, the beat it serves, what it shows, and a rough duration. The shot list is the bridge between the creative plan and the technical build. In fact it is the most important realization in this whole approach: the shot list is the plan the agent builds from, and because a human wrote it, there is no creative decision left for an AI to make at build time. This is why you do not need any kind of AI "scene generator" for a single video. You already decided what to show. The agent's job is to render those decisions, not to invent new ones.

A worked example clarifies the shape. A sales video (a VSL) is a person talking to camera with animated graphics cut in at specific moments: a hero statistic at the open, a "myth versus truth" reframe, a diagram of the core mechanism, a comparison near the buying decision, a recap, a closing credibility montage. Each of those is a named entry on the shot list with a duration and a one-line description of what it shows. That list, plus a still for each, is everything the build needs.

This guide does not re-teach how to write the blueprint itself, the hook, the structure, the persuasion architecture. For the full method, see SalesBlaster's Episode Blueprint System Guide. Here it is enough to know that the blueprint exists, that it produces a shot list, and that the shot list is what flows into the build.

When to skip this chapter. If you need a single graphic, or you already know precisely what you want, you do not need a blueprint. Your "shot list" is one line in your head. Go straight to design.

9. Conversion-leverage tiering [Studio] (one-offs may skip)

Not all graphics in a video carry equal weight, and treating them as if they do is how people burn weeks polishing things that do not matter while the things that do stay mediocre. Before you build, rank your shot list by how much each graphic affects the outcome.

A simple four-tier scheme works:

  • Tier A, carries the argument. The few graphics the whole video depends on: the opening hero visual that earns attention, the central proof, the one image that explains your core mechanism, the comparison at the decision point, the close. Build these first and polish them to broadcast grade. If any one of them is weak, nothing else compensates.
  • Tier B, supporting teaching. The graphics that explain and reinforce: process diagrams, secondary charts, timelines, supporting comparisons. Build them clean, but do not treat them as precious.
  • Tier C, illustrative flourish. Small visual beats that add texture but that the video survives without: a brief metaphor animation, a decorative transition, a minor icon. Build them last, accept a simpler treatment, and cut them under time pressure without guilt.
  • Tier D, conditional. Graphics that are only needed if a stronger asset (a real screen recording, a piece of real footage) is not available. Do not build a Tier D fallback until you know the better asset is not coming. Building a fallback for something you intend to capture is pure waste.

The tiering is governed by a principle worth stating directly, because it is the through-line of the whole system: your attention is the scarce resource, not the agent's output. An AI agent can generate graphics far faster than you can review them. So the order you build in should put your sharpest judgment where it matters most, on the Tier A graphics, while your attention is freshest. Everything in this guide is organized to protect your review bandwidth.

The persuasion layer of tiering. There is a reason the high-leverage graphics tend to be the proof and the mechanism, and it comes from how audiences buy. A viewer who has seen every competitor's claims is sophisticated and skeptical; what moves them is not another claim but fresh proof and a clear mechanism they have not seen before. So the visuals that do the heavy lifting are the ones that demonstrate rather than assert: the real number counting up, the side-by-side that makes a difference undeniable, the diagram that makes a mechanism click. Decorative visuals support; proof and mechanism visuals convert. A useful way to think about which visual type serves which moment:

What the viewer needs at this momentVisual type that serves it
Curiosity, "what am I looking at"Striking, cinematic, or sleek branded motion; a bold hero visual
Recognition, "this is my problem"Before/after contrasts, the myth-versus-truth reframe
Understanding, "this is how it works"Process flows, diagrams, step sequences, timelines
Proof, "these are the real numbers"Animated counters, scorecards, real data brought to life
Urgency, "this is real and limited"Live-feeling counters, scarcity and timing visuals

You do not need to memorize this. The point is only that graphics are persuasion, not decoration, and the ones that persuade hardest deserve your best work.


Part IV. Design it

10. The brand system (design tokens) [Both]

Before any graphic is built, you define the look once, in a single place, called the design tokens. Tokens are the centralized list of every visual constant: the colors (background, primary, accent, the green/yellow/red status colors), the fonts and their weights, the spacing and margins, and the animation feel (how snappy or smooth motion is). Every graphic reads its colors, fonts, and spacing from this one file.

There are two reasons this is non-negotiable. First, consistency: when every graphic pulls from the same tokens, your whole library shares one visual language, and an audience starts to recognize your videos before they see your logo. That recognition is brand, and brand is built by repetition of a consistent look. Second, maintainability: when you want to evolve the brand, you change one file and every graphic updates. Without tokens, a color change means editing dozens of files by hand.

Tokens are also how you handle multiple brands or sub-brands cleanly. A common pattern: a parent brand and a customer-facing sub-brand that share the same palette and motion but differ in their logo and on-screen name. You express this as one token file with a "brand" slot for the mark, so the sub-brand inherits everything and overrides only its name and logo. An agency serving many clients takes the same idea further: one token file per client, swapped per project, with the same component library underneath. The components do not change; only the tokens do. This is what lets the same studio produce on-brand graphics for any number of brands.

Because tokens gate every component, defining them is always the first build step, and nothing downstream should start until they are locked.

11. Designing the stills with Claude Design [Both]

This is where you decide what each graphic looks like, and it is the most important creative step for a non-technical operator, because it is where you exercise visual judgment without touching code.

Use Claude Design to design the still. A "still" is a single picture of a graphic at its most complete, finished moment, the destination the animation will build toward. You design it in Claude Design: the layout, the colors, the type, the card, the shadows, the composition, and the empty space around the subject. You are deciding how it looks, frozen, before anyone decides how it moves.

The division of labor is exact, and getting it right saves enormous time:

  • Claude Design designs the still. It is excellent at producing a polished static look.
  • The AI coding agent animates the still. It takes your approved picture and builds the moving, code-based version that resolves to it.

Be honest about the seam between them. Claude Design produces a picture and a starting point, not a finished animation and not engine-ready code. Treat its output as the visual specification: the agreed-upon destination. The agent then rebuilds it as an animated graphic that arrives at that destination. Do not expect to drop a Claude Design file into the engine and have it move.

Why design the still first, rather than just describing the animation in words? Two reasons. It front-loads the cheapest possible review: judging a picture takes seconds, judging a half-built animation takes minutes of scrubbing, and your review attention is the scarce resource. And it removes ambiguity: a picture is a precise target, where a verbal description leaves the agent guessing. The still becomes the contract. The animation either arrives at it or it does not.

What makes a good still:

  • Show the finished, fullest state. The number at its final value, the diagram fully built, every label present. The animation will work backward from here.
  • Bake in the brand. Use your token colors, fonts, and spacing so the still already looks like your videos.
  • Mind the subject and the environment separately. Design the subject (the meaningful thing) deliberately, and treat the background, vignette, and margins as the shared environment they are. A clean, uncluttered composition animates better than a busy one.
  • Respect the empty space. Negative space is part of the design. Do not crowd the frame; the safe margins are where the eye rests.

For a video with many graphics, you produce one still per shot-list entry. For a one-off, you produce one still. Either way, an approved still plus the graphic's intent and duration is everything the agent needs to start building.

A note on Claude Design's specific features, which evolve: check the current product documentation for exactly what it can export and how. The principle here, design the still, then animate it, holds regardless of the tool's current capabilities, and you can substitute any tool that produces a clear visual target.


Part V. Build it

12. Setting up the studio (the repo) [Studio] (one-offs use the Quickstart project instead)

A studio that produces many videos needs a single home, a repository (a "repo," just a project folder under version control). The way you organize it determines whether the system compounds or rots. The organizing principle maps directly onto the subject-versus-environment idea from Chapter 4: things that are reused across every project live in one shared place; things specific to a single video live with that video.

So the repo splits into two layers:

  • A shared layer, the durable asset, reused by every project: the design tokens, the stage (the environment wrapper and its background, safe-area, vignette, and watermark), the small set of animation helpers, and the library of reusable graphic components, along with their validation schemas.
  • A projects layer, one folder per video: that video's shot list as a small data file, its scene definitions (each one marrying a subject component to the stage with the right data and duration), any graphics built specifically for that video, and its rendered output.

The full structure:

AGENTS.md # standing instructions every agent session reads (see Appendix A)
tsconfig.json # path shortcuts so files can reference @shared and @projects
remotion.config.ts # engine configuration
design/tokens.ts # the brand system (Chapter 10)
registry.ts # index of available components
shotlist.ts # the shot list as data (the scene graph)
compositions.tsx # this video's scenes, registered for rendering
Root.tsx # top-level index aggregating every project's scenes

Notice how the three roles from the mental model map cleanly onto folders: shared/stage is the environment, shared/components plus each project's custom folder hold the subjects, and the scenes folder is where a subject gets married to the stage with real data. The architecture and the filing agree, which is what makes it easy for a person or an agent to know where anything goes.

The rule that keeps the shared library healthy: the rule of three. A new graphic is born in a project's custom folder. Only when a third project needs the same kind of graphic do you promote it to the shared library and make it configurable. Resist the urge to generalize earlier. Building a flexible, reusable component before you have proof that it will be reused is speculative work that usually guesses wrong, and it bloats the shared library with things only one video ever used. Build it specific; promote it on proof. A counter is obviously shared from day one. A graphic that explains your one particular signature concept is probably specific to you forever.

Tooling, kept deliberately simple. Do not reach for heavyweight monorepo tooling here. Start as a single engine project with the shared and projects folders and simple path shortcuts so files can find each other. One top-level index registers every scene, named by project so they do not collide. Only if a project ever needs its own separate set of dependencies should you graduate to a workspace setup. This is the simplest thing that works, and it keeps a non-technical operator (or an agent) from drowning in configuration.

This structure also positions you for the future without any work today: because the shared layer is a clean boundary, it can later be lifted into a larger company codebase wholesale if you ever need to (see Chapter 18).

13. The animation libraries [Both]

Most graphics need only the engine itself. But for specific jobs, you reach for a specialized library, and understanding what each is for, and the one rule that governs all of them, prevents the most common failures.

First, the governing idea restated, because it is the thing people get wrong: these are not separate video makers and not interchangeable with the engine. They are tools you call from inside a single graphic, and every one of them must take its timing from the engine's frame number. The engine owns the clock. A library either computes a still shape that you then reveal frame by frame, or it describes motion that you scrub to the current frame. None of them is ever allowed to run on its own internal clock, because that breaks the reproducibility that is the entire point.

What each is for:

  • The engine's own animation (the default). For the large majority of graphics, the engine's built-in tools for fading, sliding, counting, and easing are all you need. Reach for nothing else first. Adding libraries you do not need is its own form of over-engineering.
  • A data and geometry library (for charts). When a graphic is genuinely data-driven, bars, lines, curves, proportional shapes, you use a charting library for the math: it turns numbers into coordinates and shapes. You then draw the result and animate the reveal with the engine's frame. You do not use the library's own animation features, because those run on their own clock. Think of it as a calculator for "where do the shapes go," not an animator.
  • A choreography library (for complex, precisely-timed motion). When many elements must move in a tightly orchestrated sequence (this finishes, that starts a beat later, a third overlaps), a choreography library expresses that timeline cleanly. You build the timeline in a paused state and scrub it to the engine's current frame, so it becomes a description of motion that the engine plays. This maps naturally onto the engine's frame-based model.
  • A pre-made animation player (for designer-grade micro-animations). For small, polished, self-contained animations (an icon flourish, a tidy loading motion) that would be slower to hand-build, you can play a pre-designed animation file and seek it to the current frame. Best for simple, contained moments where simplicity is the whole point.
  • A 3D library (for three-dimensional scenes). For product rotations, three-dimensional diagrams, or camera fly-throughs, a 3D library renders three-dimensional scenes that the engine captures frame by frame. For simple, self-contained 3D, the plainer approach is more reliable and easier for an agent to generate correctly; for complex scenes that need to share state and compose with your other graphics, the component-based wrapper fits better. Most content does not need 3D at all, so treat this as a specialist tool.
  • A note on interaction libraries. Some popular animation libraries are designed for interactive interfaces (hover, drag, things that respond to a user). They are a weaker fit for rendered video, which is a fixed timeline with no user input. For video, the choreography library above is the stronger pairing. Mentioned only so you do not reach for the wrong tool.

Two practical techniques worth knowing, both of which an agent can apply when asked:

  • Motion along an exact path. If you need something to follow a precise curve, draw the line as an image and give it to the agent; it can trace that exact path far more accurately than it can interpret a verbal description of a curve.
  • Transparent graphics for overlays. Most graphics render as ordinary opaque video. But a graphic meant to float on top of other footage (a lower-third name tag, a corner callout over a person's face) needs a transparent background. That requires a specific render setting that preserves the transparency (a higher-quality render format with the image format set to preserve alpha). Apply it only to true overlays; using it for everything needlessly bloats render times and file sizes.

Audio, captions, and multi-format, for completeness:

  • Audio. The engine can include narration and music and mix them into the final file. For a continuous explainer you may add audio directly. For b-roll segments that a person assembles into a larger edit, audio is usually the editor's job, so you render the graphics silent.
  • Captions. Animated captions and subtitles can be produced by the engine and are worth adding for social clips, where most viewers watch without sound.
  • Multiple formats. A component can be built to adapt to widescreen, vertical, and square frames, so the same graphic can be rendered for YouTube, then re-rendered vertical for Reels and Shorts, with little extra work. Decide your formats up front and design the subject to sit comfortably in each.

Licensing, which matters commercially and which you should understand before building a business on this:

  • The charting, pre-made-animation, and most other helper libraries are free and openly licensed, including for commercial use. A leading choreography library became completely free in 2025, including all of its formerly paid plugins, with commercial use covered.
  • The engine itself is the one with a paid tier, and it is the one to pay attention to. It is free for individuals, for non-profits, for those evaluating it, and for for-profit companies with up to three employees. A for-profit company with four or more employees needs a paid company license, which starts at roughly one hundred dollars per month. Two situations push you into the paid tier: growing past three employees, and, more importantly, building a product or service that serves engine-rendered videos to your customers (for example, a tool that generates videos for clients), which falls under the usage-based commercial tier. For a solo creator rendering their own videos locally, the free tier almost certainly applies; for an agency or a software product, check the current terms against your situation. (License terms change; verify the current details before relying on them.)

14. Building the components with AI agents [Both]

This is the heart of the workflow, and it runs on three levers and one dependency rule.

The dependency rule (build the trunk before the branches). A few things must exist before any graphic can be built, and they must be built in order: the design tokens, then the small set of shared animation helpers and the stage, then the index that lets the engine find your graphics. Call this the trunk. It is a few hours of work and it is strictly sequential, because everything else depends on it.

Once the trunk exists, here is the powerful part: the individual graphics are independent of each other. A counter does not depend on a diagram, which does not depend on a chart. This is not an accident; it is the same property that lets the engine render frame 900 without rendering the 899 frames before it, applied one level up. Independence means the graphics can be built in parallel, in any order, without colliding. The build process inherits the engine's own super-power.

On the studio path, you exploit that independence with a parallel-build trick. Using version control, you can open several isolated copies of the project at once (called worktrees), run a separate agent session in each, and build several graphics simultaneously, then merge them back together with no conflicts because each touches different files. The agent can manage all of this for you. A caution, though: do not over-parallelize. You can only meaningfully review one preview at a time, so beyond two or three simultaneous graphics you are just stacking work against your own attention, which is the real bottleneck. Parallelize across families of graphics, not across individual graphics within a family, so you keep the efficiency of staying in one visual mode.

Batch your work by family, not by position in the video. Group your graphics by what kind of thing they are, all the counters together, all the charts together, all the diagrams together, rather than by where they appear in the script. Similar graphics share patterns, and an agent that just built one counter builds the second one almost for free. Building in script order would force the agent to switch between unrelated visual styles constantly and lose that reuse. (A starter set of families is in Appendix C.)

Lever 1, load the context once. At the top of your repo, keep a standing instructions file (the full template is in Appendix A) that every agent session reads automatically. It pins the non-negotiables, the frame rate and resolution, the rule that every visual value comes from the tokens, the requirement that components be self-contained, and the brand feel. Writing this once means every session inherits it, so you are not re-explaining the rules each time.

Lever 2, scope each task small and give the agent a precise target. Build one graphic per task. For each, hand the agent the approved still (the visual target), the graphic's intent, its duration, and the data it should display. A precise target plus a small scope is what produces good output. (The full prompt template is in Appendix B.)

Lever 3, iterate in the same conversation. Do not try to get a graphic perfect in one prompt. Get a simple first version on screen, then refine in small steps within the same conversation: "slow that down," "the entrance should stagger," "this color is off." The agent holds the context of the conversation, so quality compounds with each turn. Three or four turns of refinement beats one elaborate prompt. When a graphic is right, save it and start the next one fresh.

The loop, per graphic:

  1. Open an agent session (or a parallel worktree on the studio path).
  2. Give it the standing context, the still, the intent, the duration, and the data.
  3. The agent builds a first version of the component, its input-validation rule, and a small preview scene.
  4. Preview it and scrub the timeline.
  5. Refine in the same conversation until it matches the still and feels right.
  6. Approve and save.
  7. Move to the next graphic in the family.

The gates that protect quality (do not skip them): lock the tokens before building anything; build and render one graphic completely as a first proof that the whole toolchain works before fanning out; and review and approve every graphic in preview before considering it done.

The minimal build, for the one-off path. If you are making a single graphic, the loop above collapses to exactly the Quickstart in Chapter 7: one project, the engine skill loaded, one still, one conversation with the agent, one render. You do not need the trunk-and-fan-out structure, the families, or the worktrees. Use those only when you are building a studio.

15. Preview, render, and handoff [Both]

Preview is cheap; rendering is the expensive step, so preview first. The engine can play your graphic live in a browser while you scrub the timeline, with no slow rendering involved. This is where you judge timing and catch problems. Only after you approve a graphic do you run the actual render, which produces the final video file.

Match your output structure to the kind of video you are making. This is a real decision and it depends entirely on the video type:

  • A continuous animated explainer is one long graphic, so you render it as a single video file.
  • B-roll for talking-head video is a set of separate, short segments that an editor drops into the larger edit at specific moments. Render each graphic as its own file, named for the beat it serves.

In other words, the engine's output is either one reel or a folder of labeled segments, and which one you want is dictated by whether the graphics are the video or whether they punctuate a video. Most talking-head, sales, and course content is the segment case; most standalone educational explainers are the single-reel case.

Opaque versus transparent renders. Render ordinary full-frame graphics as standard opaque video. Render only true overlays (a name tag or callout that floats over other footage) with the transparency-preserving setting from Chapter 13. Do not apply the transparent setting to everything; it makes files larger and renders slower for no benefit.

Render locally until volume forces otherwise. Rendering on your own machine is fast enough for normal volumes; a short segment renders in well under a minute. Only at high, sustained publishing volume does it become worth setting up cloud rendering, which adds real infrastructure complexity. Do not reach for it early.

Output organization keeps a studio sane. Put each video's renders in a folder named for the video, name each file for the graphic, and keep a short note pairing each graphic to the beat in the script where it belongs. That note is what an editor needs to place everything correctly.

The handoff. For b-roll work, you hand an editor the rendered segments and the placement note. The editor sources any real footage, assembles the talking-head and the segments, and adds music, color, and final pacing. The programmatic pipeline gave them finished graphics they did not have to build, which moves their job from production to editorial judgment. For a continuous explainer, the rendered reel may be close to final, needing only audio and light polish. For a one-off, the rendered file is the deliverable.


Part VI. Operate it

16. The binding constraint and the anti-over-engineering rules [Both]

If you remember one operating principle, remember this: your review attention is the scarce resource, not the agent's output. Agents generate graphics faster than any human can judge them, so the entire system should be arranged to spend your judgment where it matters and to avoid building things that do not. That principle generates a short list of rules that keep the system lean.

  • Build by family and front-load the high-leverage graphics, so your sharpest review lands on the graphics that decide the outcome.
  • Do not build a studio for a single video. Use the Quickstart. Structure earns its keep at the third project, not the first.
  • Apply the rule of three before generalizing. A graphic stays specific to one video until a third video needs it. Premature reusability is wasted, mis-aimed work.
  • Do not perfect the flourish. Tier C graphics get a simple treatment and can be cut. Spend the saved time on proof and mechanism.
  • Do not build conditional fallbacks until you know the better asset is not coming.
  • Do not pull in libraries you do not need. The engine alone covers most graphics. Each extra library is surface area you have to maintain.
  • Do not add automation you do not yet need. It is tempting to build an elaborate system that auto-generates shot lists and scene plans with AI. For a single video, that is over-engineering: you already wrote the shot list, so there is no decision left to automate. (The point at which such automation does earn its place is in Chapter 18.)

The quality bar, stated plainly. The failure mode to watch for is the "generic AI look": flat, templated, characterless motion that signals low effort. The defenses against it are the ones already built into this workflow: a deliberate design system, stills designed with real visual judgment, and in-conversation refinement until a graphic feels intentional rather than default. If a graphic looks like it could belong to anyone, it is not done. A quick checklist is in Appendix F.

17. The compounding asset [Studio]

The reason to run this as a studio rather than a series of one-offs is that it gets better on three axes at once, automatically, as you use it.

  • The library grows. Your first video might require building several graphic types from scratch. Each subsequent video adds a few, driven by real need, until most of what any video wants already exists and you are configuring rather than creating.
  • The agent gets better at your style. Every approved graphic is an example of what "right" looks like for your brand. Feeding those examples back into your standing instructions makes the agent's first drafts closer to final over time.
  • Your speed climbs. As more of each video's graphics come pre-built and the agent's drafts improve, the time per video falls, from many hours to a few.

A simple discipline keeps this healthy, borrowed from a more-better-new way of thinking: first, get more out of the graphics you already have (if a component works and looks good, keep using it, even if it feels repetitive to you, because consistency is brand); second, improve the weakest link (build the next component only when you keep needing one that does not exist, or improve your instructions when the agent keeps getting one thing wrong); and only then, third, build genuinely new things or try new visual styles. Most of your library should be proven and reliable, with a small fraction reserved for experiments.

This is the same idea applied to your own operation that the system applies to content: build the asset by running the work, not instead of running it.

18. When to productize and integrate [Studio]

Two future thresholds are worth naming so you know what you are building toward and, just as importantly, what not to build yet.

When to fold the studio into a larger codebase. As long as you and your immediate collaborators are the only ones operating the pipeline, the standalone graphics repo is the right home. The trigger to integrate it into a larger company system is specific: when someone outside your core team needs hosted, self-serve access to any part of the pipeline, a team member, a partner, or a client who needs a web interface to it. At that point you need a proper application surface, and the clean shared layer you built in Chapter 12 lifts into the larger system without a rewrite. Until that trigger fires, integrating early just adds complexity you are not using.

When to add AI-driven planning. This guide deliberately keeps a human writing the shot list, because for any single video the creative decision is already made and an AI "scene generator" would only re-derive it non-deterministically. The point at which automated planning earns its place is when you have several videos queued whose shot lists you genuinely do not want to write by hand, and your component library is stable enough that an AI can reliably map a narrative to it. That is the seed of an AI content agent that takes a video idea and produces the graphics with no human in the loop, the eventual end state of this whole system inside a larger AI operating model. It is a real destination, but it is the last thing you build, not the first. Reaching for it before then is the textbook version of using an exciting future capability to avoid the unglamorous present work of just shipping the video in front of you.


Part VII. Reference

Appendix A. The AGENTS.md template [Both]

Place this file at the root of your project and name it AGENTS.md, the cross-agent standard. Every agent session reads it automatically, Claude Code included: it reads AGENTS.md whenever no CLAUDE.md exists, so one file serves every agent tool and you maintain a single source of truth. Fill the brackets to match your tokens.

# Graphics build context

## Non-negotiables
- Resolution and frame rate: [1920x1080], [30] frames per second. Time is measured in frames, never seconds, in code.
- Every color, font, spacing, and easing value comes from src/shared/design/tokens.ts. No hardcoded visual values anywhere.
- Prefer spring-based motion over plain linear motion for entrances. Use full-frame layout for backgrounds and timed sub-sequences for internal timing.
- Every component is a self-contained, pure function of the current frame and its inputs. No side effects, no data fetching while rendering, no dependence on other components for timing.
- Every component ships with: an input-validation schema, and a small preview scene with realistic sample data, registered so it appears in the preview.

## Brand
- Visual identity: [dark background, high contrast, clean and modern, engineering-grade, your accent color]. Not playful, not cartoonish, not generic.
- If a graphic looks like default, templated, characterless AI motion, it is wrong. Rebuild it.

## How we work
- One graphic per task. Build the simplest version that shows the core idea first, get it on screen, then we refine together in this same conversation. Do not over-build the first pass.
- The provided still is the target: the finished graphic should resolve to it.

Appendix B. The component build prompt template [Both]

Clone this per graphic and fill the brackets. Attach the still.

Build one animated graphic. Follow the standing context file.

NAME: [graphic name]
DURATION: [N] seconds
WHAT IT SHOWS: [one or two sentences describing the graphic and its purpose]
THE DATA TO DISPLAY: [the real numbers, labels, or text]
THE TARGET LOOK: [attached still]. The finished graphic should resolve to this image.

DELIVERABLES:
1. The component, reading all visual values from the tokens, self-contained.
2. Its input-validation schema.
3. A small preview scene using the real data above, registered so it shows in the preview.

FIRST PASS: build the simplest version that puts the core idea on screen (the count, the bars,
the diagram, the reveal). Then I will refine timing and polish with you in this conversation.

Appendix C. Component family taxonomy and starter catalog [Studio]

Group graphics into families for batched building. A starter set, with the recurring patterns most content needs:

  • Counters: a large animated number with a label; supports a secondary comparison figure. (Stats, results, "before and after" numbers.)
  • Text and cards: headline reveals, key-point lists that stagger in, the myth-versus-truth reframe (strike the old belief, reveal the new one), offer and call-to-action cards, titles.
  • Charts: bar comparisons, funnels, curves, proportional figures, health or progress meters. (Data brought to life.)
  • Flows and diagrams: process pipelines with data flowing through stages, two-node give-and-get loops, self-improving cycle diagrams, step sequences.
  • Interface mockups: a clean fake calendar filling with events, a research-to-result sequence, a sorting or ranking animation, a set of variants with a winner highlighted.
  • Comparisons: side-by-side this-versus-that, a deliberately tangled "doing it the hard way" diagram, a wide-net-versus-focused metaphor.
  • Micro and brand: small flourishes, badges and shields, scroll and pointer cues, the logo sting, intro and outro cards.

You will not build all of these for any one video. Build only what a video's shot list calls for, and promote a graphic into this shared catalog only when a third project needs it.

Appendix D. The shot list schema [Studio]

Express the shot list as a small data file so the build is precise and reproducible. Conceptually, each entry carries: a unique name, the component or custom graphic it maps to, the family it belongs to, its duration in frames, the data and labels it displays, and its leverage tier. A validation layer should confirm that each entry names a real component and supplies the inputs that component requires, catching mistakes before any rendering happens.

If you ever extend toward automated planning (Chapter 18), this same schema is the contract an AI would produce, and the validation layer is what keeps an AI from inventing a graphic that does not exist. Until then, you write this file by hand, because the shot list is a creative decision you have already made.

Appendix E. Commands cheat sheet [Both]

# Create a new project
npx create-video@latest [project-name]        # choose: blank template, TypeScript, install

# Teach your agent the engine (run inside the project)
npx skills add remotion-dev/skills

# Start the AI coding agent
claude

# Preview live in the browser (scrub the timeline, no slow render)
npm run dev

# Render the final file
npx remotion render [composition-name] out/[file].mp4

# Render a transparent overlay (preserves alpha for floating graphics)
npx remotion render [composition-name] out/[file].mov --codec=prores --prores-profile=4444 --image-format=png

# Other formats
npx remotion render [composition-name] out/[file].webm --codec=vp8
npx remotion render [composition-name] out/[frames] --sequence

If a command is unfamiliar, paste it to the agent and ask it to run it and explain what it does. You do not need to memorize these.

Appendix F. Quality checklist [Both]

Before calling a graphic done:

  • Does it match the approved still?
  • Do all colors, fonts, and spacing come from the brand tokens (does it look like your videos)?
  • Is the timing right when scrubbed, not too fast to read, not so slow it drags?
  • Does it feel intentional, or does it look like default, templated motion?
  • For an overlay, is the background actually transparent?
  • For proof and mechanism graphics specifically: is it as clear and credible as you can make it, since these carry the persuasion?

Appendix G. Troubleshooting [Both]

  • A command says a tool "is not recognized." A prerequisite is missing or not installed correctly. The most common is the runtime (Node.js). Reinstall it and reopen the terminal.
  • The agent writes code that does not behave like the engine should. Reload the engine instructions (npx skills add remotion-dev/skills) and restart the agent so the rules are in context.
  • A graphic does not appear in the preview. It was not registered in the index. Ask the agent to confirm the graphic is registered and the name matches.
  • A render fails. Often a version mismatch among the engine's packages. Ask the agent to check that all engine packages are on the same version.
  • The output looks generic. Go back to the still and the brand tokens, and refine in conversation. Generic output is almost always under-direction, not a limitation of the tools.
  • In doubt, ask the agent. Paste the error or describe the problem; the agent can usually diagnose and fix it. This is the advantage of a workflow where the agent does the technical work.

Appendix H. Glossary [Both]

  • Programmatic graphics: animated video graphics produced by code rather than by editing clips on a timeline.
  • The engine (Remotion): the framework that turns code into video by drawing each frame and stitching the frames into a file. It owns time.
  • Frame: one image in the video. At 30 frames per second, frame 30 is the one-second mark. Every animation reads the current frame to know what to draw.
  • Deterministic: the same inputs always produce the same output, exactly. The property that distinguishes this approach from AI video generators.
  • Component: a self-contained, reusable graphic (a counter, a chart, a diagram). The "subject."
  • Subject: the meaningful thing on screen and any effect attached to it (its shadow, its glow). Lives in the component.
  • Stage / environment: the shared wrapper holding everything constant across graphics: background, vignette, watermark, safe margins. Built once, reused everywhere.
  • Design tokens: the single file defining all brand visual constants (colors, fonts, spacing, motion feel). Every graphic reads from it.
  • Still: a single picture of a graphic's finished, fullest moment, designed before animation. The target the animation builds toward.
  • Blueprint: the plan for one video, including who it is for and what each scene shows.
  • Shot list: the ordered inventory of every graphic a video needs, with names, durations, and descriptions. The bridge from plan to build. Also called the scene graph.
  • Scene: one graphic placed in the stage with its real data and duration, ready to render.
  • AI coding agent (Claude Code): the tool that writes and runs the code for you from plain-English instructions. Your builder.
  • Standing instructions (AGENTS.md): the file every agent session reads automatically, pinning the rules so you do not re-explain them.
  • Family: a group of similar graphics (all counters, all charts) built together for efficiency.
  • Rule of three: keep a graphic specific to one video until a third video needs it, then promote it to the shared library.
  • Render: producing the final video file from a graphic.
  • Opaque vs transparent render: a normal full-frame video versus a graphic with a see-through background for floating over other footage.
  • Worktree: an isolated copy of the project that lets you build several graphics in parallel without conflicts (studio path).
  • The binding constraint: your review attention, which is scarcer than the agent's ability to generate, and which the whole system is arranged to protect.

This guide codifies a repeatable method. The tools it names will evolve; the principles, deterministic graphics directed by a human and built by an agent, the subject-and-environment split, the brand tokens, the still-then-animate workflow, the family-batched build, and the protection of your review attention, are the durable core. Verify current product and license details against official documentation before relying on them.

On this page

How to produce broadcast-quality animated graphics for any video, using AI agents and code, without a motion designer and without learning to animate.How to use this guideTable of contentsOpeningPart I. Understand it1. What this is, and what it is not [Both]2. Why it works and why it matters [Both]3. Who it is for, and everything you can make [Both]4. The mental model: four actors and one rule [Both]Part II. Decide your path5. Two paths: one-off vs studio [Both]6. Before you begin (prerequisites) [Both]7. Quickstart: one animation, fast [One-off]Part III. Plan it8. The episode blueprint [Studio] (one-offs may skip)9. Conversion-leverage tiering [Studio] (one-offs may skip)Part IV. Design it10. The brand system (design tokens) [Both]11. Designing the stills with Claude Design [Both]Part V. Build it12. Setting up the studio (the repo) [Studio] (one-offs use the Quickstart project instead)13. The animation libraries [Both]14. Building the components with AI agents [Both]15. Preview, render, and handoff [Both]Part VI. Operate it16. The binding constraint and the anti-over-engineering rules [Both]17. The compounding asset [Studio]18. When to productize and integrate [Studio]Part VII. ReferenceAppendix A. The AGENTS.md template [Both]Appendix B. The component build prompt template [Both]Appendix C. Component family taxonomy and starter catalog [Studio]Appendix D. The shot list schema [Studio]Appendix E. Commands cheat sheet [Both]Appendix F. Quality checklist [Both]Appendix G. Troubleshooting [Both]Appendix H. Glossary [Both]