Skip to main content
Back to Articles

Ever-Improving Conversations

Clarity First has supported voice interviews from the beginning. This year we rebuilt the conversation itself on real-time voice. Here is what changed, and why now was the moment.

22 July 2026Tim Farmer

Before an organisation spends real money on a change, it needs to understand the problem it is actually solving. Most don't. Leaders can end up in something of an echo chamber, and the true voice of the business, across every level and every function, is slow and costly to gather. That discovery stage is where momentum stalls, risks go unmanaged, and opportunities quietly slip away. Clarity First is built to fix that. It is a power tool that augments discovery: it reaches everyone who matters in a fraction of the usual time, and traces every finding back to what people actually said.

At the centre of it is the interview. Participants can sit down with a consultant in person or online, or take a structured interview on their own time with our AI interviewer. This post is about that last option, and a change we have made to how it works.

Why now was the right time for real-time voice

Real-time voice has been an obvious destination for AI interviewing for a while. The question was never whether we would build it, only when.

Clarity First has supported voice from the beginning. Participants could speak using push-to-talk or type, within the same interview, and behind the scenes the interviewer agent (Socrates) already understood context, followed the conversation, knew when to probe, and adapted its questions to what it was hearing. The interviewing capability was there. What wasn't there was the conversation itself: every exchange still meant pressing a button, speaking, releasing it, and waiting for a reply. It worked well, but it never quite faded into the background the way a good conversation should.

Until recently, the part that wasn't ready was the real-time conversation itself. Earlier models were impressive in a demo, but the behaviour that matters in an interview wasn't there yet: turn-taking felt awkward, interruptions weren't handled naturally, and a conversation could lose its rhythm. Over the past several months that has changed at a pace that genuinely surprised us. Today's models handle interruptions gracefully, tell the difference between a pause for thought and the end of one, hold context across the whole conversation, and recover naturally when both sides talk at once. For the first time the technology mostly disappears, and that was the moment it was ready for Clarity First.

The interview hasn't changed. The conversation has.

It is worth being clear that this is not a new interviewing method. 'Socrates' has not suddenly learned how to interview; it already knew how. It follows the same discovery framework, asks the same thoughtful follow-ups, and holds the same standards of neutrality and evidence. What has changed is how the interview feels. Participants simply talk. They can pause to think, interrupt, correct themselves, or overlap with the interviewer the way they would with another person, and the interviewer handles those moments while keeping the full thread of the conversation. Text is still there whenever it suits better, for accessibility, a shared office, or plain preference, but for most people the spoken interview is now the default.

Better for participants, and better for discovery

The point isn't that speaking is faster. It's that speaking is more natural. People explain things differently out loud: they remember an example halfway through an answer, change direction, tell a story. Those detours are often where the richest insight sits, and they rarely survive a text box. There is another effect too. With no one to impress and no judgement in the room, people often open up more than they would with a consultant. Practically, a participant can pause and pick the interview up later from any device, and if a connection drops the conversation reconnects where it left off. Whatever is said still flows through the same analysis pipeline and produces the same traceable evidence that stands behind every Clarity First report.

Built at the right moment

The milestone that mattered wasn't that the models got smarter. It was that they got conversational enough for the technology itself to get out of the way. We never wanted the interview to feel like operating a piece of software; we wanted it to feel like talking to a thoughtful person. Only recently could it, and that is why now was the right time.

Try it

The best way to understand the change is to be interviewed by it. Take a short sample interview, about five minutes, and stop whenever you like. You will feel how quickly it stops being software and starts being a good first meeting. Try the voice demo.

Ready to try it yourself?

Take a short sample interview, about five minutes. You can stop whenever you like, no account needed.

Try the voice demo