A 15-minute video from a script, five dots and a plush moose

Blog post #80


OpenAI held its DevDay keynote on September 29. Within hours I had a 15-minute video about it on YouTube, and a Swedish version rendering behind it. No camera, no avatar, no video editor. A script, my cloned voice and a folder of HTML.

It was the most fun I have had building anything this year, and I learned a lot along the way. Here is the short version.

What changed since last log

  • The presenter is gone. I no longer pay for an avatar service. The on-screen company is now five small animated characters, drawn as simple shapes, that float in, blink, hop and point at things. They are my own sketches in the spirit of the “dots” OpenAI showed on stage.
  • The video is code. I moved from stitching clips together with command-line jobs to HyperFrames, where a video is an HTML page with a timeline. Every scene is a file I can read, change and rebuild.
  • A plush moose joined. ChatGPT made me a plush moose called Algot with a full set of animations. He waves hello in the intro, looks sad in the chapter about Europe and jumps when good news arrives.

What shipped

  • The English video, Super Intelligence Is Here: Everything OpenAI Just Announced. 14 minutes 52 seconds, eleven chapters, explained for someone who has never used an AI tool.
  • A Swedish version with translated voice and translated on-screen text. The full cut is 17 minutes 28 seconds. The one for YouTube is a 14 minute 40 second cut, for reasons below.
  • A reusable kit. The characters, the colors, the chapter transitions, a small library for stepping through Algot’s sprite sheet, and an authoring guide that a helper can follow.
  • The intro I wanted. Three ticks and a ping while a four-petal logo builds itself, one petal per tick, and the play arrow lands on the ping.

What’s working

One reference chapter, then parallel helpers. I built chapter 1 myself until it looked right, wrote a one-page guide, and let four agents build chapters 2 to 11 in parallel. They matched the style well, and they found a real bug for me: the gradients on the characters shared names across chapters, so a blue dot could turn yellow. One fix in the shared file and every chapter was right.

Sync by word, not by guess. My first cut changed slides a second before I finished the sentence. The fix was to transcribe my own voice track on my own machine, which needs no account and no internet, and then run every animation through a small function that moves it from a rough time to the exact word time. When I rewrote a paragraph, the old scenes froze during the new sentences and I added fresh scenes in exact time.

A voice file per chapter. Because each chapter is its own audio file, rewriting one paragraph costs a handful of credits instead of regenerating a quarter of an hour of speech. The newest ElevenLabs model made my voice noticeably nicer to listen to, in both Swedish and English.

Two reviewers for the translation. I wrote the Swedish from the meaning of each sentence, not word by word. Then an agent and GPT, through the API, each reviewed it. They caught different things. One of them noticed that “flyttats fram” can mean “postponed” in Swedish, when the keynote said “moved earlier”. Both flagged the word for “plan” as an English habit when the Swedish word is “abonnemang”. Using two reviewers was well worth it.

Real sources, no guessing. I pulled the keynote captions, checked claims against several sources, and wrote down where each logo came from. The one claim I could not confirm in a source, that X’s bots are live in Europe, comes from my own use, and I said so.

What’s unclear or broken

  • Custom thumbnails are off until my channel is verified next year. YouTube’s own suggested frames will do until then.
  • The 15-minute limit for unverified channels. The English video fits with a few seconds to spare. I uploaded the full Swedish cut anyway, and YouTube deleted it on the spot. I cut two chapters (the workspace one and the one for companies), sped the voice up by three percent, which nobody will notice, and got it down to 14:40. The full version waits until my channel is verified.

Decisions made

  • Positive by default. The only sad moment in the video is Europe waiting for dots, and even that ends on good news, because GPT-6.1 Sol is available here.
  • Everything on screen is redrawn. No clips from the keynote, no screenshots. Logos are used only to identify the products, with their sources written down.
  • No music. The tick-tick-ping stinger was enough.
  • Swedish is its own video on the same channel, not a dubbed track.
  • Cut, don’t squeeze. Removing whole chapters kept the story clean, and the audio stayed untouched apart from the small speed-up. Re-rendering took about ten minutes.

Tooling & process

  • Secrets can stall. My password manager’s command line waited for an approval dialog I could not see, and long jobs sat on “authorization timeout”. A local, no-account fallback for the transcription kept the work moving.
  • Never delete before the replacement exists. I removed five audio files before the new ones were generated, and the generation then failed. I recovered them by cutting the chapters out of the assembled track. Now a file is overwritten only after the new one succeeds.
  • Check after you copy. Copying a project folder overwrote the Swedish audio with the English one. A quick length check caught it.
  • Long jobs need their own process. My upload waiter was killed by a ten-minute tool limit. I moved it to a detached process with a lock file, so it runs once and survives.

— Stefan