Skip to content
All work
AI Product · Semantic VideoOpen source

VidzAI

Turn audio or text into a planned visual story — then render it as MP4.

An audio/text → visual-story platform that understands meaning, plans scenes, matches assets and renders captioned MP4s with Remotion.

Problem

Turning a script or narration into a video usually means slides, stock clips in upload order, or a model that invents footage. The result often ignores what is actually being said.

Context

Open-source platform: React 19 + TypeScript + Vite on the front, Express on the API, Remotion 4 for render. A deterministic ontology — not an LLM — turns English, Hindi or Odia source text into the same scene plan. Unconfigured providers are labelled unavailable instead of faked.

Architecture

Audio or text is detected for language, understood as concepts, split into meaning-based scenes, planned visually, matched to assets, captioned, then rendered. Remotion draws the timeline — it does not decide the story.

VidzAI architecture
  1. Input
    • Audio upload
    • Pasted script
    • Optional assets
  2. Language
    • Detect en / hi / or
    • Keep source language
  3. Meaning
    • Ontology
    • Topics
    • Entities
    • Intent
  4. Plan
    • Scene boundaries
    • Visual strategy
    • Asset relevance
  5. Render
    • Captions
    • Remotion
    • MP4 + audio

Engineering decisions

Deterministic ontology, not a default LLM
The same idea in English, Hindi or Odia produces the same visual concept. Captions and narration stay in the source language.
Remotion renders the plan — it decides nothing
Language, scene cuts and asset choice happen before render. Remotion draws motion graphics, images and captions from the timeline.
Honest capability registry
GET /api/capabilities and the dashboard show live status. Speech-to-text is a mock until a provider is configured. Image analysis, AI image/video and stock search report not available instead of returning placeholders.
Relevance threshold, no random fill
User assets match on file name, description and tags. Scores below 0.45 are rejected. There is no Math.random() in visual selection.
Original audio is never regenerated
Uploaded narration is muxed into the MP4. A render without an audio stream is treated as failed.

Who it is for

Designed for people who already have a script, lesson or narration and need a structured video — not a blank-canvas editor.

  • Content creators and YouTubers turning writing into scenes
  • Marketing and education teams building explainers
  • News and media turning stories into visual timelines
  • Developers who want a foundation for semantic video generation

What is actually available

The dashboard lists generation capabilities from server configuration. Nothing is labelled AI-generated unless a configured provider produced it.

  • Available without keys: audio processing, language detection, ontology, scene planning, user-asset matching, Remotion motion graphics, captions and render
  • Mock today: speech-to-text — paste the script for accurate captions, or use a demo clip with a stored transcript
  • Not configured: image content analysis, AI image/video generation, stock search — set provider env vars on the server to enable them

05 images · open slider

Lessons

  • If a provider is not configured, say so on the dashboard. Do not return a fake image or a silent MP4.
  • Scene cuts should follow meaning. A timer does not know when the topic changed.