
An AI platform that turns a written script into a finished YouTube video: narration, animated characters, captions and metadata, with the rendering handled by its own service.

Doodrio is an AI platform we designed and built end to end. A creator pastes in a script, chooses how it should sound and look, and gets back a publish-ready YouTube video: narration, animated characters, word-synced captions, cinematic motion and the title, description and tags to go with it.
It is built for advice channels (self-improvement, psychology, motivation, stoicism and the like) where the value sits in the writing rather than in footage. The product is equally clear about what it is not for: anything that needs real video, such as cooking, travel, sport or product reviews.
Turning a script into a watchable video normally means a production team or a stack of disconnected tools: one for voiceover, another for visuals, another for captions, another for thumbnails and metadata. Every handoff is manual, and the whole sequence has to be repeated for every single video.
Creators writing advice content had the sharpest version of this problem. The script was the entire product, but the production work wrapped around it took longer than the writing, and outsourcing that work was rarely worth it for a channel still finding its audience.


Take a script all the way to a publishable video without a person touching the production, while leaving enough control that everything doesn't come out looking the same.
Three services, each with one job.
A Next.js app is where the creator writes or pastes a script, picks the voice, captions and background, and reviews the scenes before committing to a render. A Laravel API sits behind it, orchestrating the OpenAI calls for script rewriting and speech, handling accounts and OAuth sign-in, and running billing, with a Filament panel for admin and support.
Rendering is deliberately kept apart. A Node and TypeScript worker picks jobs off a queue and assembles the video on its own, outside the web request cycle. A render takes around half an hour, so it was never going to sit inside an HTTP request, and keeping it separate means render load never slows down the app a user is clicking on.


From script to finished video:
Around the pipeline:
The marketing site and the app are Next.js and React with MUI. The API is Laravel, with JWT authentication, Socialite for OAuth, the OpenAI PHP client, Stripe for billing, and a Filament panel for admin and support work.
The renderer is its own Node and TypeScript service. It does the heavy lifting, so it runs as a set of independently managed production processes with its own queue, and can be scaled or restarted without touching the app.
That split is the point of the project from an engineering view: a slow, resource-hungry job is isolated from the thing users are actually clicking on, so heavy render traffic cannot degrade the app, and a failure on one side does not take the other down with it.


Doodrio is in open beta with subscribers on paid plans and an active support channel. Videos are being rendered and published through it, and the pipeline has held up well enough that the work now is about output quality rather than keeping it running.
For us it is the reference project for infrastructure work: several services in production together, a queue doing genuinely heavy lifting, third-party APIs for language and speech, and billing alongside all of it, not a prototype that would fall over under real use.