How AI works out that a question was aimed at you
Sergei Skrylkov, Founder, Trippi · 2026-09-07 · 7 min read · facts checked against the product on 2026-09-20
A live copilot that reacted to every question in a call would be unusable — on a panel, most questions are not yours. Detection is therefore two separate problems: is this line a request at all, and was it addressed to you. They are solved in different places, for reasons that are mostly about cost and latency.
Recognised speech does not arrive as sentences
The first surprise for anyone building this: there are no reliable question marks. Speech recognition emits a stream of words with speaker labels and unstable punctuation, corrected as more audio arrives. A question may land as “so how would you handle that” with no mark at all, or split across two segments.
So detection cannot lean on punctuation. It leans on opening words, on imperative forms, and on who has been speaking.
Stage one: the cheap filter that runs on your machine
Most of what is said in an interview is not a request: acknowledgements, thinking noises, the interviewer narrating their own screen. A local check drops those before anything leaves the browser. It costs nothing, sends nothing, and exists for two reasons — money and quiet.
It is deliberately generous. It throws away the obviously empty, not the doubtful, because a missed question is expensive and an extra model call is cheap.
Stage two: the model, and the only question it is asked
What survives goes to a model with the recent lines, the participant names and your own name. The judgement it makes is narrower than people expect: was the closing turn addressed to this person? “Karl, do you want to take this one?” and “and what would you have done?” are grammatically similar and land on opposite sides of that line.
This is also why your name matters. In a call where nobody says your name and the platform does not label you, the model is working from position alone — it will be right most of the time and wrong more often than it needs to be.
Stage three: the gate on the way back
An answer arriving is not the same as a card appearing. Below a confidence threshold, nothing is shown at all. The same question twice in a minute is merged rather than shown twice. With automatic answers off, only what you typed gets through.
Those three rules are why the window stays quiet for most of a round, and quiet is the feature — a copilot that talks constantly gets closed before the interesting question arrives.
What detection still gets wrong
- Rhetorical questions. “Who wouldn’t want that?” is a question in form only. Context usually settles it; sometimes it does not.
- Names mangled by recognition. If your name is heard as another word, the strongest signal disappears silently.
- Long multi-part questions. A turn with three questions in it produces one card, usually answering the last part best.
- Panels where nobody uses names. Four voices and no labels is the hardest case there is.
Why not just wait to be asked
A hotkey would remove the whole problem — and it would also mean pressing a key while somebody watches you think. Every manual trigger is attention spent at the exact moment you have none. That is the trade: automatic detection is sometimes wrong, and a manual one is always visible.
The reasonable middle is both, which is what Cue does: automatic cards, plus a field where you can ask outright when it stayed quiet.
Trippi Cue
An interview copilot that works inside the call: it hears the question addressed to you, suggests a short answer, and translates the conversation underneath. No bot joins, and it is a browser extension rather than an application.
Answers while the interview is still running.
Trippi Cue is in the Chrome Web Store and starting is free. One click on the icon when the call begins.
