From video to a citation
you can check yourself.
Retrieval over video, built for market research rather than for meetings. This page is the pipeline stage by stage, with what it gets right and where it fails.
Eight stages between the file and the answer.
Everything on this page is built and tested. None of it is a plan.
- 1
Transcription with separated voices
Each speaker is kept apart from the others, with a timestamp on every word.
- 2
A role and a name for each voice
The moderator is identified by the fact that she asks the questions. Participants get their names when the session shows them, the way a video call does with its labels.
- 3
Played stimuli, marked as such
When a commercial or an audio clip is played inside the session, it is marked as a played stimulus instead of being credited to a participant. Without this, a question about opinion would be answered in the voice of the advertisement.
- 4
Repeats collapsed
When the same commercial is played several times, the copies are collapsed so that they do not bury the reactions to it.
- 5
The image, every fifteen seconds
Three frames per stretch are read for the scene, the people, the brands, and the on-screen text word for word. It is what answers a question whose evidence is only visual: a label, a pack, a price on a slide.
- 6
Hybrid search
Meaning and exact words at once, with Spanish matched with or without accents, reordered by a relevance model and cut off at a relevance floor.
- 7
A fair share between sessions
With fifty videos loaded, every relevant session contributes material. The ones that did not discuss the topic are declared as such, because that is a finding too.
- 8
Citations that cannot be invented
Each fragment reaches the model as a document with citations turned on, so a citation can only point at a fragment that exists. If the answer is cut off halfway, it is rejected as incomplete rather than shown to you.
99.3% of citations verbatim, over 37 questions on 11 videos.
The test ran on 25 September 2026 across three collections of real sessions. Every one of the 275 citations was compared against a second, independent transcription produced by a different engine, and 273 matched.
The automatic comparison cleared 94.9%. The 14 cases it flagged were read by hand: 12 were differences between the two transcriptions, which is to say the sentence was in fact said. The other 2 are the 0.7%.
The failures the same test found.
A video mixing Spanish and Portuguese had a Portuguese passage translated into Spanish by the transcription. A citation from there is faithful in meaning and not literal.
Heavily edited video and purely visual questions are the weak spot. The visual collection scored worst of the three.
Image analysis can fail on individual stretches. The dialogue survives, and half the time for those stretches is returned.
Processing time is not promised. The only end-to-end measurement so far is a synthetic 46-second clip, which says nothing useful about a two-hour session.
Your sessions are yours, and the plumbing is what says so.
- Every account sees only its own projects, enforced in the database rather than in the interface.
- Videos live in a private bucket and play through signed links that expire quickly.
- The transcript is deleted from the transcription service once it has been processed.
- The image analysis flags sensitive content stretch by stretch. Past a threshold the video is blocked, it is not indexed, and its hours are not returned.
Ask your focus groups anything. Every answer cites the exact minute.
Bring five hours of sessions or forty. Finding the moment and quoting it is mechanical work; what it means for the brand is still yours to decide.
Want to try it on your own sessions?
It is in beta with invited researchers. Tell us what you are working on.
Ask about the beta