Improving AI outputs is actually a lot trickier than it looks. We’ve gone through dozens, if not hundreds of cycles of iteration with little to no progress to make up for the work, spanning across several AI tools we’ve been concocting up to improve our internal workflows.
After months of work, we’ve finally hit our first 4/5 when it comes to quality output while AI consistency is now growing right along side it.
(Note: this is one of our spreadsheets that allow us to measure all of our tools’ successes across the board. The “Quality Outputs” bucket is useful in measuring and tracking AI output across a project’s lifecycle)
One thing that we’re learning is that understanding absolutely matters. In context, this means that if you don’t know how an output is achieved nor how it’s measured, it’s extremely difficult to get AI to match the quality of work that you want it to produce.
People think that AI is a one stop shop. But ultimately, what you put in is what you get out.
Here are a few tactics that we’ve put together to get better AI outputs (note that a lot of this aligns to how normal product work is done):
The feedback loop is everything. Without it, the work falls flat, nothing worthwhile gets worked on, and the team spins in circles. Do short, fast cycles. Get feedback from external users and internal team members. Use it to elevate the work.
Measure and define what matters. If you don’t have this worked out, you won’t have a way to determine success. We typically like spreadsheets, measuring particular groups of desired functionality on a 1-5 scale.
Know what great output looks like. The people normally doing the work should have a surmountable amount of feedback coming from them.
Understand the labor required to get a great output. Talk to those doing the work. Have them break down their processes, and how they manually go from point A to point B. As an implementor, you have to have a decent level of understanding to orchestrate.
It’s been a crazy few months, but we’ll be testing this workflow more with some of our next tools as we get to more 4/5’s and hopefully 5/5’s. Very exciting work ahead of us.
Yaaaas! We’re learning quite a bit along the way. There are many layers to the build and development understanding needed to accomplish what we’re after.
Feedback guides the work itself. Imagine sketching out a bridge, only yourself, a pencil, and a piece of paper. You might draw up something that looks like a bridge, but there’s no way to know that it works.
Gravity is the feedback mechanism for will it hold, cars (or people) are the feedback mechanism to know that it works, and other designers are the mechanism to know that it looks good.
You can’t build a good bridge without leveraging feedback in some way. It’s exactly why software that replicate cars and weights on a bridge is SUPER valuable, and a team is required to help you catch things that you’ll miss as you’re heads down in the work itself.
The thing that we sometimes forget, especially in the product realm, is that everything is connected to people. Without loads of human feedback (or maybe even synthetic someday), you can’t build a great product.
When building AI tools and increasing the quality of AI output, we’re having to re-envision what some of that feedback looks like. And, even though it is pretty similar to what we’re already doing, the whole AI part has thrown us for a bit of a loop. But, we’re getting our footing with boots laced up (but maybe missing pants).
Great read @ben! Really tells the story of the iterative loop it takes to get these tools right. We expect that Output Consistency score to go up with repetition. Excited to start showing off this Call Rubric tool to people!