Open source · bring your own inference

The agent harness for video. Sees the screen. Hears the words. Checks the facts.

A free, open-source agent that turns any video into structured, evidence-linked data, right from your terminal.

From a video to data you can actually use

Most tools hand you a transcript and call it data. Lirovo gives you the facts: the claims, decisions and numbers inside the video, structured, scored, traceable to the exact second they came from, and checked against their sources.

01Ingest

Drop in any video

Paste a link or upload a file. A YouTube video, a Zoom or Teams recording, a webinar, an MP4. Lirovo fetches it and handles the messy parts for you.

See supported sources
02Capture

It hears every word and reads every screen

Lirovo transcribes the audio and reads what is on screen at the same time: slides, charts, captions, on-screen text. The half that is spoken and the half that is shown, both captured.

See how capture works
03Timeline

It links both onto one timeline

Spoken words and on-screen moments are aligned second by second, then connected into a live graph of the people, topics, numbers, claims and decisions inside the video.

See the graph
04Schema

You get clean data, shaped to your schema

Tell Lirovo the fields you want: names, claims, decisions, numbers, topics. It returns typed JSON with a confidence score on every value, ready for a spreadsheet, a database or your app.

See an example
05Evidence

Every value carries its proof

Each value links back to the exact second it came from, and whether it was said out loud or shown on screen. Lirovo flags where the two disagree and checks claims against their sources, so you can verify every finding yourself.

How evidence works
Run it standalone, or inside your agent:Terminal appPluginBring your own inferenceOpen source
claude · ~/lirovo
❯Try “build a timeline of every number he cites”
⏵⏵ auto mode on(shift+tab to cycle)·lirovoMCP connected
Standalone, or in your agent

Run it in your terminal, or inside your agent

Lirovo runs as its own terminal app, or as a plugin inside the agent harness you already use. Hand it a link, get back typed JSON with a citation for every value, all on your own inference.

  • extract() hand it a video and a schema
  • get_job() get back typed JSON
  • get_graph() walk the knowledge graph
  • get_evidence() cite the exact moment
  • deliver() send it wherever you work
one commandlirovo extract <url>

Why developers reach for Lirovo

Minutes
From video to typed JSON

Most videos finish end to end in minutes, not hours. The heavy media work, scene detection and frame dedup, runs before any model call, so your inference is spent only where it matters.

Learn more
100%
Of values are traceable

Every extracted value links back to the exact second it came from, with its modality recorded. Claims are checked against their sources, so nothing is a black box.

Learn more
1 pipeline
Audio, visual, and reasoning

Transcription, on-screen reading, and reasoning run as one local pipeline on the inference you already use. Nothing to wire up, and your video never leaves your machine.

Learn more

Frequently asked questions

Still have questions? We are happy to help you scope your use case.

Talk to us

Start turning video into data

Get early access