eDiscovery has become sophisticated at handling documents. The technology stack for finding relevant text in millions of emails, contracts, and records is mature: TAR workflows, predictive coding, concept clustering, privilege log automation. These tools have meaningfully reduced the cost of document review over the past 15 years.
Video discovery is at an earlier stage. The standard approach for most litigation matters involving video evidence is still linear human review: a paralegal or junior associate watches recordings, takes notes, and flags potentially relevant moments for attorney review. In matters involving hundreds of hours of deposition footage, surveillance recordings, body camera footage, corporate training videos, or broadcast archive material, this approach is economically painful and practically risky. The pain is easy to quantify. The risk — what gets missed because it isn't worth the billable hours to find — is harder to see but arguably more significant.
The Math of Manual Video Review
Consider a hypothetical employment discrimination matter. The opposing party has produced 320 hours of video from corporate training sessions, town halls, and recorded HR interviews conducted over a three-year period. The question is whether specific statements were made about promotion criteria, demographic patterns, or management practices that are relevant to the claims.
At a 1:1 ratio — one hour of review time per hour of footage — that's 320 attorney or paralegal hours just for first-pass review. At realistic billing rates, that's a significant discovery cost for a single evidence category. In practice, the first-pass ratio often runs higher: reviewers stop, rewind, take notes, flag clips for export. 1.5:1 or 2:1 is common for content that requires careful attention. Deposition footage runs even higher because every exchange may be relevant and the phrasing of questions matters.
In larger matters — class actions, regulatory investigations, commercial litigation involving operational video — the scope can be orders of magnitude larger. At some point, the economics make full review impractical regardless of budget. Something gets cut. What gets cut is usually the material that's deemed lower-probability relevant. That's a reasonable judgment call under resource pressure, but it's a judgment call made on the basis of titles and rough timestamps, not on the basis of actual content. Relevant material gets missed not because anyone was negligent, but because nobody could afford to look.
What Indexed Video Changes
When video content is indexed — transcripts with speaker diarization, OCR of on-screen text, semantic embeddings across all modalities — the review workflow changes structurally. Instead of watching content to find relevant moments, reviewers run targeted searches and watch results.
For the training session example: a search for semantic concepts related to promotion criteria, combined with speaker-attributed transcript search for specific executives' names, returns a set of time-stamped clips across the 320-hour corpus. Instead of watching 320 hours, reviewers watch the returned clips — potentially hours, not weeks. They then exercise judgment on the retrieved results: relevance determination, privilege review, proportionality decisions. The judgment work is the same; the time spent watching non-relevant material is dramatically reduced.
The secondary benefit is defensibility. When producing a privilege log or preparing a discovery certification, an indexed review has a more complete paper trail. The search queries run, the results retrieved, the review decisions made: these are documented systematically. A manual review produces notes; an indexed review produces a structured record. In contested discovery proceedings, that distinction can matter.
Specific Video Categories in Legal Contexts
Different video evidence categories have different indexing priorities.
Deposition recordings are primarily speech content. Speaker diarization matters more than visual indexing, because the value is in who said what and when. Accurate diarization that correctly attributes speech to the witness versus counsel is essential; misattributed speech in a deposition transcript produces misleading search results. WER on deposition audio tends to be low — controlled environment, single speaker at a time, no background noise — so ASR accuracy is generally good. The main technical challenge is multi-day depositions where audio characteristics shift between sessions.
Surveillance and body camera footage is primarily visual. There may be little or no speech content, but visual recognition of people, objects, locations, and events is critical. Frame-level scene analysis and visual embedding search are the dominant useful modalities. OCR is secondary — incidental scene text (license plates, signage) may be relevant but is lower density than in broadcast content.
Corporate video — recorded meetings, training sessions, town halls — tends to be mixed: significant speech content plus structured on-screen material (presentation slides, shared screens). Both ASR and frame OCR are important. The presenter often refers to slide content without reading it verbatim, so the index needs to capture both the spoken content and the displayed text separately. A search for a specific figure or policy statement may match on the slide text rather than the audio.
The Chain of Custody and Authenticity Consideration
Any legal use of indexed video content requires attention to chain of custody for the source material. Indexing is a derivative process — it reads from source files and produces metadata. The source files themselves need to be preserved and hash-verified per standard eDiscovery protocols (EDRM guidelines or equivalent). Meshora's index is a search layer over the source material, not a replacement for it. Every search result links back to a specific timecode offset in a specific source file; that source file needs to be in your document management system with intact provenance tracking.
We're not saying video indexing removes the need for standard discovery protocols — those exist for reasons that extend beyond search efficiency, and an indexed corpus still needs to comply with preservation obligations, collection procedures, and production formatting requirements. What indexing changes is the efficiency of the review step in the discovery workflow, not the legal and procedural framework around it.
There's also an authentication consideration when relying on transcript content from ASR. If a transcript will be used to characterize what a person said in a video, the ASR output needs to be treated as a working tool, not as a certified transcript. WER means there will be errors. Before using specific language from an ASR transcript in a filing, you verify against the source video. The transcript is a finding tool; the video is the evidence.
Where the Economics Actually Shift
The meaningful change isn't in small matters where video volumes are modest. If you have 12 hours of deposition footage, the cost difference between manual review and indexed search is real but not transformative.
The meaningful change is in matters where video volume was previously a practical barrier to full review. When 300 hours of footage can be searched in the same time it previously took to watch 20 hours of it, matter teams make different decisions about what to include in their review scope. Content that would have been written off as too expensive to review fully becomes reviewable. That changes what gets found, and occasionally what that means for the outcome of the matter.
A litigation support team managing a matter with a large video corpus — say, 600 hours of recorded employee testimony from a workplace investigation — faces a different set of options when search latency is seconds rather than weeks. Keyword searches for specific names and phrases across the full corpus take minutes to run. Semantic searches for conceptual relevance take minutes to run. The decision about whether to search for a specific concept becomes "should we do this?" rather than "can we afford to do this?" That's not a minor shift — it changes the practical threshold for what gets investigated and what doesn't.
Video evidence in litigation is going to increase. Recorded calls, video depositions, body cameras, security systems, corporate communications moving to video formats: the proportion of case-relevant material captured in video is growing. The discovery infrastructure for handling it efficiently is worth building now, not after the matter scope has already been scoped down to what's affordable to review manually.