Blog
EngineeringJune 12, 20263 min readLirovo team

From Conference Video to Training Data: Structured Annotation for Multimodal Models

Teams training video-language models need structured temporal annotation, not bounding boxes. How conference video becomes training data.

The teams building the next generation of video-language models (VLMs) need structured training data. Not bounding boxes. Not weak labels from web scraping. Temporal grounding, speaker attribution, visual-audio alignment, and verifiable claims with citations.

Today, that training data is expensive, manual, and slow. Teams either pay Scale AI $10-50 per hour of video for generic annotation (designed for bounding boxes, not temporal understanding), or they build their own extraction pipeline. Neither option produces the right format.

What training pipelines actually consume

The standard for video training datasets in production systems is HuggingFace's WebDataset and Parquet format. (HuggingFace) LeRobotDataset v3.0, used by teams training embodied AI and robotics models, uses Apache Parquet for tabular metadata (frame ID, timestamp, entity labels, claims), MP4 for video, and JSONL for rich per-sample annotation.

  • Parquet table: frame-level features, temporal markers, entity mentions, claim linkage
  • JSONL sidecars: one line per frame with full evidence spans and grounding
  • Frame indices: timestamped references to the video stream
  • Metadata JSON: episode/clip boundaries, duration, speaker roles

What matters: everything is temporally grounded. A model learns not just "this is a claim" but "this claim appears from t=120 to t=145 with audio span [120.3-144.8] and supporting evidence from frames [frame_542, frame_551]."

What Lirovo already produces

When you upload a conference keynote, product demo, earnings call, or technical talk to Lirovo, the extraction pipeline outputs:

  • Timestamped transcript with word-level alignments
  • Knowledge graph nodes (claim, entity, product, action) each with t_start and t_end
  • Evidence spans linking each claim back to the frame, audio segment, and spoken context it came from
  • Scene-boundary keyframes with vision vectors from Gemini
  • Temporal cross-references between transcript, graph, and frames

This is exactly what a team training a video-understanding model needs. The hard part (extraction, temporal grounding, multimodal alignment) is already done.

A 90-minute keynote becomes 2,000 labeled frames, a knowledge graph with temporal evidence, and a structured transcript in under 20 minutes, at roughly $0.40.

The domain where this works

Not all video training data is equal. Lirovo is built for domains where speech is dense, declarative, and visually grounded:

  • Conferences and keynotes: technical talks, product announcements, investment theses, research presentations
  • Earnings calls and investor meetings: financial claims, business updates, guidance, Q&A
  • Product demos: feature walkthroughs, workflow explanations, use-case narratives
  • Technical webinars and instruction: how-to content, troubleshooting, educational lectures

These are precisely the domains where VLM training data is most deficient and most valuable. Vision-language models trained only on static images miss the temporal and narrative structure. Models trained on weak video labels (from web scraping) miss the grounding. Lirovo fills that gap.

vs. Scale AI for video training data

Manual annotation (Scale AI)

Cost: $10-50 per hour of video

Automated (Lirovo)

Cost: ~$0.40 per 90-min video

Manual annotation (Scale AI)

Turnaround: days to weeks

Automated (Lirovo)

Turnaround: minutes (real-time for short clips)

Manual annotation (Scale AI)

Temporal precision: frame-level only

Automated (Lirovo)

Temporal precision: frame + audio segment + transcript span, linked

Manual annotation (Scale AI)

Output: bounding boxes, classifications

Automated (Lirovo)

Output: knowledge graph, evidence spans, multimodal linkage

Manual annotation (Scale AI)

Domain: generic (images, bounding boxes)

Automated (Lirovo)

Domain: specialist (video intelligence, speech-grounded claims)

The training data vertical

The market opportunity is real. The global AI training dataset market is projected to grow from $2.82B in 2024 to $9.58B by 2029 (CAGR 27.7%), with multimodal data as the fastest-growing segment at 31.1% CAGR. (MarketsandMarkets)

Within that, the annotation-tools market alone spans $1.69B (2025) growing to $14.26B by 2034. Critically, automated annotation already accounts for 58% of market share versus 42% manual. (Fortune Business Insights)

Scale AI proved the business model: $870M revenue in 2024 grew to $2B in 2025. But Scale AI dominates generic annotation (bounding boxes, classification). Video training data requiring temporal grounding and multimodal structure is still a manual, expensive, slow process.

Lirovo can own the specialist vertical: structured temporal annotation of video content where speech, visual context, and temporal alignment matter. That is a much smaller market than "all annotation," but it is a market where the unit economics are dramatically better and the switching cost is high.

Export formats coming soon

We are building export options so your Lirovo extractions become ready-to-train datasets:

  • HuggingFace export: Parquet table + JSONL sidecars, directly uploadable to HF Hub
  • WebDataset shards: tar-compressed samples for distributed training pipelines
  • Custom JSONL: frame-aligned structured data with full evidence and temporal markers

No re-annotation needed. The temporal grounding is already there.

Try it on your own video

Export training data from your videos with timestamped frames, structured claims, and temporal evidence.