Capture

It hears every word and reads every screen.

Two streams run side by side. Lirovo transcribes the spoken audio and reads what is on screen at the same time, then aligns both onto one timeline. The half that is spoken and the half that is shown, both captured.

Two halves of the same moment

Most tools only do the audio half. They hand you a transcript and stop there. Lirovo captures both halves and keeps them in step.

Hears

The audio, transcribed speaker by speaker

The speech is transcribed and attributed to who said it, so a quote is never a free-floating line of text. Every word lands at the second it was spoken, which is what lets it line up with what was on screen at that same second.

Reads

The screen, read frame by frame

A vision model reads what is actually on screen: the slides, the charts, the captions and chyrons, the dashboards, the code. The figure a presenter points to but never says out loud is still captured, because Lirovo saw it.

What reading the screen catches

01
Slides and charts

The numbers on a slide, the labels on a chart, the axis a value sits on. The figures a presenter points to but never reads out loud.

02
Captions and chyrons

Lower-thirds, news tickers, name plates, and burned-in captions. The on-screen text that names who is speaking and what they represent.

03
Dashboards

A metric on a screen-share, a status in a console, a row in a table. The state of a system as it was shown in the moment.

04
Code on screen

A function in an editor, a command in a terminal, a config being edited. The exact text shown, not a paraphrase of it.

Only what changed

It studies the frames that matter, not every frame

Reading every single frame would be slow and wasteful, and most frames are near-identical to the one before. So Lirovo first normalizes the video, detects where the scene changes, and removes near-duplicate frames. What is left is the set of frames that actually changed.

Only those frames are read by the vision model, while the audio is transcribed alongside them. That keeps capture fast and cheap, with no redundant work and nothing on screen missed.

Mistral CEO Arthur Mensch
analysis ready
Mistral CEO Arthur Mensch
Mistral CEO Arthur Mensch
CNBC · 45:59
Timeline0:00 / 45:59
Arthur
Arjun
Events
0:0011:2922:5934:2945:59
Compute & sovereignty
AI economics
Semiconductors
Vibe & agents
Cybersecurity
AGI & physical world
ContradictionDecisionClaimclick to scrub
Speakers2
Arthur Mensch
92%
Arjun Karpal
4%
b_rollFRAME · 0:29

A stylized shot of the Eiffel Tower in Paris at dusk.

Objects detected · 2
Eiffel Towercityscape
read by VLM · gemini-3.1-flash-lite

Aligned on one timeline

The spoken words and the on-screen moments are lined up second by second onto a single timeline. A claim made out loud at one second sits next to the chart that was on screen at that same second, so the two can confirm each other.

That timeline is what later becomes the knowledge graph: the people, topics, numbers, and decisions inside the video, connected to the exact moments they came from.

Every model call in this pipeline runs on your own inference, so you use the best model for each step, transcription, vision, and reasoning, and you are never locked to one provider.

See the graph it builds

Capture is the first half of the story. Next, both streams are linked onto one timeline and connected into a graph you can query.