We've been testing and learning and testing again · 1two.one
All posts
20 August 2026·Alex·3 min read·Share

Where We've Been: Testing, and Then Testing Again

It's been a couple of weeks. Not because there was nothing to say — because we've been heads down doing the kind of work that matters more than almost anything else we build.

It's been a couple of weeks since the last post here. Not because there was nothing to say, just because we've been heads down doing the kind of work that doesn't make for exciting content in the moment but matters more than almost anything else we build.

Testing. Then testing again. Then testing the thing we just fixed, to make sure fixing it didn't break something else.

Real conversations…

Building an AI system that reads a supervision conversation and understands what actually happened is a genuinely hard problem, and we didn't want to find that out the hard way with real providers and real data.

So before anything went near a real user, we built a testing framework of realistic supervision scenarios. Not simplified examples. Full, messy, true-to-life conversations, the kind that meander, that circle back, that bury the most important thing in a sentence that starts with "it's probably nothing."

Each scenario was written to be genuinely difficult. Not difficult in an artificial, gotcha way, difficult in the way real supervision conversations are difficult, because people rarely say exactly what they mean the first time, and the most important thing in the room is often the thing said last, almost as an afterthought.

Testing, comparing, testing again

We didn't run these once and move on. Some scenarios have been run repeatedly, against different configurations, different models, different versions of the underlying instructions, comparing results each time to see what improved and what didn't.

The testing process has been ever evolving. As the picture has become clearer, we've been able to test not just our own instructions but the underlying technology choices behind them, and at times that's revealed genuinely stark contrasts. The same conversation, processed two different ways, can produce meaningfully different pictures of what actually happened in the room, and of how much genuine evidence of good practice was really there to find. That's exactly the kind of finding this whole process was designed to find.

When a test revealed the system had missed something, that became a fix. When a fix was made, the same test ran again, and others alongside it, to make sure the improvement was real and hadn't quietly broken something else. Then more scenarios were added, harder ones, longer ones, ones designed specifically to probe whatever the last round of testing had shown us we needed to understand better.

This is still going. It doesn't stop. Every improvement opens up the next question worth testing.

What this has actually taught us

A few things worth sharing without giving away the mechanics of how the system works underneath.

The most important insight isn't always the first thing said in a session. Sometimes it's the thing raised right at the end, introduced with "it's probably nothing" or "I don't want to make a big deal of this." Good supervision AI has to weight significance correctly, not just chronologically.

People rarely say exactly what they mean the first time. Minimising language, hedging, changing the subject, these aren't noise to filter out. They're often exactly the signal worth paying attention to.

Two things can be true at once. A worker can be right about a concern and have gone about raising it the wrong way. Good analysis holds both, rather than collapsing into one simple verdict.

And genuinely good practice needs to be as visible in a record as anything that needs addressing. A system that only ever surfaces problems isn't giving a manager the full picture of the person in front of them.

Why we're telling you this

None of this shows up as a feature you'll click on. It's not a button or a dashboard. It's the invisible work underneath everything else, the reason a summary reads the way it does, the reason the right thing gets flagged for a manager's attention and the wrong thing doesn't create noise.

We think that work deserves to be visible, even if the outputs of it aren't the kind of thing that makes an exciting screenshot. Trust in a system like this has to be earned by demonstrating the rigour behind it, not just by describing the features on the surface.

Alongside all of this testing, we've also been doing significant work on the technical foundations underneath, data handling, security, and how the platform is architected. That's the subject of the next post.

For now, know this: the quiet work has been happening. It's ongoing and it's making the product genuinely better, one difficult conversation at a time.

Found this useful?

Share it with your network — help another registered manager save an hour today.

Share on LinkedIn