Four indexing layers.
One search interface.
Every spoken word. Every on-screen character. Every visual scene. Every entity connection. Meshora fuses them into a single knowledge graph that returns timestamped results in under 50 milliseconds.
Five stages from raw file to searchable knowledge.
Ingest
Upload directly or point to your S3, GCS, or Azure Blob. Meshora reads file manifests and queues processing automatically.
Transcribe
Speech-to-text with speaker diarization across 13 languages. 94% accuracy. Every spoken word timestamped to the frame.
Detect
14 scene-level signals per frame. OCR text extraction from chyrons, tickers, slides. Visual emotion classification. Object and person recognition.
Graph
Named entities extracted and linked across clips: people, places, products, organizations. The same expert appearing in 400 clips gets connected automatically.
Search
Index written to our vector store. Natural language queries return ranked segment results with timecodes, confidence scores, and transcript excerpts.
Four ways to find any moment.
Natural language
Ask in plain English: "find the moment they discuss the product recall." Meshora interprets intent and returns the best-matching clips ranked by confidence.
"spokesperson mentions safety concerns"
Exact transcript
Quote verbatim: Meshora finds every occurrence of that exact phrase in your transcripts, timestamped to the second. Useful for compliance and legal review.
"recall of all units manufactured before"
Visual concept
Describe what the scene looks like: "close-up shot of documents on a table" or "speaker at podium, formal setting." Returns clips matching the visual description.
"outdoor press conference, news microphones"
On-screen text
Search chyrons, ticker text, slide content, overlaid captions. Any text that appears in frame — not just what was spoken — is fully indexed and searchable.
chyron: "BREAKING NEWS"
Three data streams. One search result.
Audio, vision, and text don't live in separate buckets — they describe the same moment from different angles. Meshora's fusion layer embeds all three into a shared vector space so a query matches against all signals simultaneously.
- Speech embeddings from 94%-accuracy ASR with speaker attribution
- Visual embeddings from per-frame scene classifiers and object detectors
- Text embeddings from OCR-extracted on-screen text
- Fused at index time — no runtime merging overhead on queries
- Not a transcription-only tool — speech is one of four signals, not the whole product
Numbers from real archive indexing jobs.
Based on internal testing across regional broadcast archives and corporate training libraries. Your mileage varies by content type.
Connects to your existing storage. No migration required.
pip install meshora — production-ready client
npm install @meshora/sdk — TypeScript types included
Try Meshora on your archive.
14-day trial. 10 free hours. No credit card. Your archive, indexed and searchable before tomorrow.