The 9-box grid plots performance against potential to give a leadership team a shared, portfolio-level view of talent. Its weak spot is the potential axis — a subjective, once-a-year opinion. Prove keeps the grid but scores that axis from measured behavior — Initiative, Applied Grit, and Learnability — so each placement rests on proof, not a show of hands.
| The 9-box grid | Prove | |
|---|---|---|
| What it measures | Performance on one axis, potential on the other, plotted into nine cells. | Three observed behaviors: Initiative, Applied Grit, and Learnability. |
| How “potential” is judged | A subjective read — the manager or the room forms an opinion. | Scored from real behavior over a defined window, not a show of hands. |
| Bias exposure | Open to recency and halo bias; the loudest voice often anchors the cell. | Anchored to what a person actually did, which narrows the room for bias. |
| Cadence | Typically a once-a-year snapshot at a talent review. | A repeatable cycle you can run whenever you need a fresh read. |
| Output | A position on a grid and a shared label for the conversation. | A behavioral score you can point to and defend with evidence. |
Nine named cells give a leadership team one vocabulary for talent, so the conversation is about the same thing.
Plotting the whole team at once surfaces gaps, over-reliance, and bench depth you would miss person by person.
The grid is a structured prompt for a discussion that might otherwise never happen — and structure beats no structure.
Performance is observable, but “potential” is usually a boardroom opinion — prone to recency and halo bias, and to whoever speaks with the most confidence.
A once-a-year placement freezes a person in a cell, long after the behavior that put them there has changed.
When someone asks why they landed in a given box, “we felt so” is a weak answer — and a risky one to act on.
This is the same trap that shows up whenever the read is a judgment rather than a measurement — the theme of it’s the metric. And it is why we prefer scoring behavior over inferring it from a questionnaire, as covered in behavioral vs. personality tests.
This is not either-or. Keep the 9-box for what it does well — the shared language and the portfolio view of your whole team. Then, instead of judging the potential axis by a show of hands, score it with behavioral evidence. You run the same talent review, but every placement is backed by what people actually did. The grid stays; the guesswork goes. If you want to know how much of your current read is proof versus gut feel, the Certainty Diagnostic is a fast place to start — and identifying high-potential employees breaks down the behaviors that belong on that axis.
The 9-box is only as good as the judgment behind its potential axis. To keep a talent review honest, anchor that axis to evidence rather than adjectives. In practice, that means working through a checklist:
The 9-box grid is a talent-review tool that plots employees on two axes — performance (usually the x-axis) against potential (the y-axis) — sorting people into nine cells, from low performance and low potential up to high on both. Its strength is a shared language and a portfolio view of the whole team; its weakness is that the potential axis is a subjective judgment.
In most 9-box reviews, potential is not measured at all — it is judged. A manager or a leadership group forms an opinion and places the person in a cell, often by a show of hands. That makes the axis vulnerable to recency bias, halo effects, and the most confident voice in the room. Scoring the axis from observed behavior instead is what turns the guess into evidence.
The performance axis is as reliable as your performance data. The potential axis is only as reliable as the judgment behind it — and unstructured judgment is prone to bias and drift, especially in a once-a-year snapshot. The 9-box is a genuinely useful discussion framework; it becomes more reliable when the potential axis is anchored to behavioral evidence rather than opinion.
The 9-box is a framework for sorting talent into a grid; the placement on the potential axis is a subjective read. Prove is a measurement — it scores Initiative, Applied Grit, and Learnability from real behavior over a defined window. Put simply, the 9-box gives you the grid and the shared language; Prove gives you a defensible number for the axis the grid guesses at.
Yes — that is the recommended way. Keep the 9-box for its shared language and portfolio view, but score the potential axis with Prove's behavioral evidence instead of a show of hands. You run the same talent review, but each placement is backed by what people actually did, so the conversation moves from opinion to proof.
Book a call to see how Prove turns the guess on your 9-box into measured behavior.