Hypit: Clone Any Viral Video With AI Agents — a Word-Anchored Workflow, Not a One-Off Render
Most “AI video” tools hand you a finished MP4 and call it a day. Hypit takes a different bet: give the agent — Claude Code, Codex, OpenClaw — a language and runtime for video, and let it build a workflow you can edit and re-run. Drop in a viral clip and the agent doesn’t just transcribe it; it clones the whole production as an editable composition: footage, word-level captions, B-roll, and effects, all anchored to words instead of seconds.
Word-anchored, not timestamp-anchored
The core idea is deceptively simple. Everything in a Hypit composition is tied to words rather than timecodes. Rewrite a line and the caption timing re-flows itself. Swap the narrator and the captions stay put. Translate the whole thing to Spanish and the three-stage reveal still lands on the same beats. That’s the mechanism behind the headline pitch — 1 command, 100 variants — because the parts that stay the same never get touched, and only the parts that change get regenerated.
It’s an agent Skill
You install Hypit into your coding agent:
npx skills add hypit-ai/hypit -g
Then you talk to it like any other skill:
/hypit Clone this video: /path/to/video.mp4, and replace the ranking
content with a comparison of Hypit with other AI video products.
Or skip the reference entirely — “Make a ranking video that puts Hypit in S tier” — and the agent writes the workflow from scratch, requests whatever credentials the render needs, generates the assets, and assembles the final composition. The Skill is compatible with Claude Code, Codex, and OpenClaw, and the repo notes integration with OpenAgents Workspace for orchestrating agents across devices.
The receipts
What makes the project credible is that it publishes real examples with itemized production notes and costs:
- “GOAT DEBATE” — a 20-second football tier list (Ronaldo in D, Messi in S). Two Seedance 2 Mini A-rolls, a 2K portrait and ten B-rolls from GPT Image 2, WhisperX word alignment, a sound-synced ranking board, color-box karaoke captions, rendered concurrently across 64 headless Chromium processes. Three clones included. Total: $1.15.
- “DAILY CREATINE” — an 18-second split-screen podcast ad with speaker-aware captions and a product-handoff moment. Three ad verticals from one format. Total: $1.07.
- “NICE RIDE” — a 26-second street interview with head-tracked speaker-colored captions (driven by Google Video Intelligence + YOLOv8 face boxes) and an emoji reveal board. Total: $1.09.
Generation models are optional throughout — a workflow can compile captions, motion graphics, and code-rendered visuals into a finished video without calling any generation API, which is how these land around a dollar.
Genuinely open, genuinely pointed at ads
Hypit is open source with no seat pricing, per-render fees, or forced watermark; you pay only the model service you choose (their hosted HypiHub, your own API, or local models). The use cases it leads with are unambiguous: clone a winning ad from the Meta Ad Library and ship 50 hook variants the same day, turn any Reel into a template, spin a new TikTok Shop video per SKU.
That power cuts both ways. A tool explicitly designed to mass-clone other people’s viral videos — faces, voices, and all — raises real questions about originality, likeness consent, and platform spam policies. If you’re using it for your own footage and brand, it’s a remarkably efficient pipeline. If you’re cloning strangers’ videos at scale, that’s a decision worth making deliberately, not by default.
For anyone building agent-driven content pipelines, though, Hypit is one of the more interesting primitives to appear lately: it treats video as editable, re-runnable code that an agent can author — not a black-box render.
Repo: github.com/hypit-ai/hypit · Demo: hypit.ai