Video in the works
AstraTech Video Studio
A rendered AstraTech promo, and how a story video is put together.

About the project
An AI helper that turns narration and recordings into branded AstraTech videos. A teammate opens it in Claude Code or Codex, AI coding assistants, and just asks for a video.
AstraTech needs a lot of video: promos, course stories, student stories. Video Studio makes them from code. A teammate opens it in Claude Code or Codex and asks for a video in plain words, and the agent builds it from Remotion templates.
Two templates work today: a 21-second brand promo, and a vertical cartoon story of any length where the scene changes on each spoken line and subtitles follow every word.
How it works
Record
Narrate the story, in Hindi or English.
Transcribe
whisper.cpp writes down every word with its timing. A corrections file fixes mistakes without moving the timing.
Build scenes
Scripts turn the story into scenes: a cut on each spoken line, camera moves and chapter cards.
Render
Remotion renders in chunks, so a 15-minute video doesn't need 14 GB of disk at once.
Hand it over
It's packaged as one agent skill that a teammate installs just by asking.
Under the hood
- Templates
- Remotion, React 19 and TypeScript.
- Audio
- ffmpeg and whisper.cpp.
- Scene scripts
- Python.
- Delivery
- A Claude Code skill that also works on Windows.
What I did
- Two video templates in code: a 21-second brand promo and a vertical cartoon story of any length
- A brand kit: colours, an animated logo and background
- Story videos with scene changes on each spoken line, camera moves, chapter cards and word-by-word subtitles
- A local pipeline: Hindi speech-to-text with word timings, a corrections file, scripts that turn a story into scenes, and chunked rendering for long videos
- Tested on a 15-minute Hindi story with 58 scenes and 10 characters
- Packaged it as one agent skill that a non-technical teammate can install by asking the AI, on Mac or Windows
What was new
- My first Remotion project: videos written in React
- Tying every cut to a spoken word instead of a fixed time
- Packaging a team workflow as an agent skill
Problems I hit, and how I fixed them
One 15-minute vertical render needed about 14 GB of temporary frames.
Fix Render in chunks, about 1.5 GB at a time.
Hindi transcription needed about 100 fixes per 15 minutes.
Fix A corrections file, with cuts tied to words so fixes never shift the timing.
Transcription leaves out filler words like "umm".
Fix Find them in the gaps between words that still contain sound.
Cut-out characters pasted over photos looked bad, and the team rejected it.
Fix Full illustrated scenes instead.
Screenshots
Screenshots in the works
Real screens from AstraTech Video Studio go here.