Drop in any video
Paste a link or upload a file. A YouTube video, a Zoom or Teams recording, a webinar, an MP4. Lirovo fetches it and handles the messy parts for you.
See supported sourcesOpen source · bring your own inference
A free, open-source agent that turns any video into structured, evidence-linked data, right from your terminal.
Most tools hand you a transcript and call it data. Lirovo gives you the facts: the claims, decisions and numbers inside the video, structured, scored, traceable to the exact second they came from, and checked against their sources.
Paste a link or upload a file. A YouTube video, a Zoom or Teams recording, a webinar, an MP4. Lirovo fetches it and handles the messy parts for you.
See supported sourcesLirovo transcribes the audio and reads what is on screen at the same time: slides, charts, captions, on-screen text. The half that is spoken and the half that is shown, both captured.
See how capture worksSpoken words and on-screen moments are aligned second by second, then connected into a live graph of the people, topics, numbers, claims and decisions inside the video.
See the graphTell Lirovo the fields you want: names, claims, decisions, numbers, topics. It returns typed JSON with a confidence score on every value, ready for a spreadsheet, a database or your app.
See an exampleEach value links back to the exact second it came from, and whether it was said out loud or shown on screen. Lirovo flags where the two disagree and checks claims against their sources, so you can verify every finding yourself.
How evidence worksLirovo runs as its own terminal app, or as a plugin inside the agent harness you already use. Hand it a link, get back typed JSON with a citation for every value, all on your own inference.
extract() hand it a video and a schemaget_job() get back typed JSONget_graph() walk the knowledge graphget_evidence() cite the exact momentdeliver() send it wherever you workMost videos finish end to end in minutes, not hours. The heavy media work, scene detection and frame dedup, runs before any model call, so your inference is spent only where it matters.
Learn moreEvery extracted value links back to the exact second it came from, with its modality recorded. Claims are checked against their sources, so nothing is a black box.
Learn moreTranscription, on-screen reading, and reasoning run as one local pipeline on the inference you already use. Nothing to wire up, and your video never leaves your machine.
Learn moreStill have questions? We are happy to help you scope your use case.
Talk to us