We’ve used UX metrics to track people for years. What they did, how they felt, how fast they got through a task. We bucketed them into Attitudinal, behavioral, performance.
Taste and delight…attitudinal.
Habits and interaction loops…behavioral.
Speed, accuracy, completion…performance.
All of it pointed at the human. When we started building Radar, where the system writes the content, it’s clear that a new area of metrics had emerged in the connection between bot and human. @ben and I talk through some of what we have been learning.
Intelligence? How well the system reasons, aligns to what you meant, and communicates it back. That’s not a human metric..it’s more of a system one. And you can’t improve what you don’t measure.
This is why I’m excited about the feedback loop we built into Radar. On the surface it’s simple. Every section of an article gets a thumbs up or thumbs down and a quick note. It looks like basic quality control.
It’s more than that. Every rating is a signal about whether the system reasoned and communicated the way we wanted. This is somewhat subjective to taste. That’s the start of an Intelligence metric. Capture enough of them, feed them back into the prompt through an agent, and you’re closing a loop on something the old metrics never tracked.
We kept the feedback binary on purpose. We don’t fully know yet what a good Intelligence metric looks like, so we’re capturing simply now and can change how we capture it later.
Here’s where we’re exploring, and where I’d love to hear from you.
-
Is a thumbs up enough signal? Or do you need something richer to actually tune a system?
-
How do you turn “this reads better” into something a model can optimize against?
-
What does a good Intelligence metric even look like when the thing you’re measuring is reasoning, not clicks?
We’re early on this. But it feels like the content feedback loop is pointing at a whole quadrant of measurement that’s going to matter more as more of the work gets automated.
Curious what the rest of you are seeing in your product work. If you’re running feedback loops on AI-generated work, what are you capturing, and is it telling you anything useful yet?


