How do you benchmark a team's AI ability?
You benchmark a team's AI ability by measuring the same real-work signals consistently across people and over time: which tasks each person can improve with AI, whether the gains are verified, and how capability trends month to month. Consistent evidence, not a one-off test, is what makes a benchmark meaningful.
A benchmark is only useful if it is comparable. That means measuring capability the same way for everyone, from real work rather than self-report, so differences reflect ability and not who rated themselves generously.
It also has to move. A single snapshot tells you where a team was; a living benchmark tells you where it is going. Tracked over time, it shows which teams are pulling ahead and where support will pay off.
Yes, if you measure it the same way from real work. Comparisons built on quizzes or self-ratings are unreliable, because the inputs differ. Comparisons built on verified work hold up.