It hears every word and reads every screen.
Two streams run side by side. Lirovo transcribes the spoken audio and reads what is on screen at the same time, then aligns both onto one timeline. The half that is spoken and the half that is shown, both captured.
Two halves of the same moment
Most tools only do the audio half. They hand you a transcript and stop there. Lirovo captures both halves and keeps them in step.
The audio, transcribed speaker by speaker
The speech is transcribed and attributed to who said it, so a quote is never a free-floating line of text. Every word lands at the second it was spoken, which is what lets it line up with what was on screen at that same second.
The screen, read frame by frame
A vision model reads what is actually on screen: the slides, the charts, the captions and chyrons, the dashboards, the code. The figure a presenter points to but never says out loud is still captured, because Lirovo saw it.
What reading the screen catches
The numbers on a slide, the labels on a chart, the axis a value sits on. The figures a presenter points to but never reads out loud.
Lower-thirds, news tickers, name plates, and burned-in captions. The on-screen text that names who is speaking and what they represent.
A metric on a screen-share, a status in a console, a row in a table. The state of a system as it was shown in the moment.
A function in an editor, a command in a terminal, a config being edited. The exact text shown, not a paraphrase of it.
It studies the frames that matter, not every frame
Reading every single frame would be slow and wasteful, and most frames are near-identical to the one before. So Lirovo first normalizes the video, detects where the scene changes, and removes near-duplicate frames. What is left is the set of frames that actually changed.
Only those frames are read by the vision model, while the audio is transcribed alongside them. That keeps capture fast and cheap, with no redundant work and nothing on screen missed.

A stylized shot of the Eiffel Tower in Paris at dusk.
Aligned on one timeline
The spoken words and the on-screen moments are lined up second by second onto a single timeline. A claim made out loud at one second sits next to the chart that was on screen at that same second, so the two can confirm each other.
That timeline is what later becomes the knowledge graph: the people, topics, numbers, and decisions inside the video, connected to the exact moments they came from.
Every model call in this pipeline runs on your own inference, so you use the best model for each step, transcription, vision, and reasoning, and you are never locked to one provider.
See the graph it builds
Capture is the first half of the story. Next, both streams are linked onto one timeline and connected into a graph you can query.
