Percipio logo

Validating It Worked

Measuring the Capability Against the Strategy That Initiated It

Matt
Matt
Validating It Worked
Validating It Worked

Most organizations that get the sequencing right, strategy first, then capability, then AI where it earns its place, eventually run into a different problem. Someone in leadership asks a simple question: did it work? And the honest answer, more often than anyone would like, is "we're not sure."

That's not for lack of data. Capabilities built on a solid strategic foundation still generate enrollment numbers, completion rates, satisfaction scores, and usage dashboards for the tools that came with them. The problem is that none of those numbers answer the question being asked. They describe activity. They don't say whether the organization can now reliably do the thing the strategy required, or whether that capability moved the business outcome it was funded to produce.

Getting from "we built the capability" to "it worked" requires a different kind of discipline than building the capability in the first place. It's worth being just as precise about validation as about the initial capability definition.

Two Different Questions, Not One

"Did it work" is actually two questions wearing one sentence, and most measurement efforts quietly answer only the easier of the two.

The first question is whether the capability itself got built: can the organization reliably produce the outcome without depending on the two or three people who happened to be strongest in the room, and does the process, tooling, and incentive structure around the work actually support it. The second question is whether that capability, now built, moved the strategic result it was meant to serve: the new market entered, the operating model that got faster, the regulatory exposure that got smaller.

The initiative to develop a capability can often pass the first test and fail the second. Teams may become genuinely capable at something the strategy or business no longer needs, especially if the initiative ran long enough for the business context to shift underneath it. Treating these two questions as one question is how a report full of green checkmarks ends up next to a strategic result that never moved. Picture a team that spends a year building a sharp analytics capability. They are well-staffed, well-documented and genuinely reliable, but it’s for a market that the business has already decided to exit. The first test passes cleanly. The second one fails just as cleanly, and no dashboard built on leading indicators alone would catch it.  

With the two questions separated, the next challenge is finding measures that can actually answer each one.

Build a Chain of Evidence, not a Single Metric

No single number carries enough weight to prove a capability initiative worked. What holds up under scrutiny is a chain of evidence across three tiers, each one a necessary check on the others.

  • Leading indicators show whether the new behavior is actually happening, for example, is the new process being followed, is the tool being opened, is the decision being made the new way instead of the old way. These often show up more quickly, but they are also the easiest to game. They become vanity metrics. They can be useful as an early warning system but aren’t sufficient as proof of anything on their own.
  • Capability indicators show whether the organization can produce the outcome reliably and without relying on specific individuals: consistency of output quality across people and teams, how the work holds up when a key person is out or moves on, and how much rework or escalation the new way of working still generates. This is the tier most measurement plans skip, and it's the one that distinguishes organizational capability from individual skill.
  • Strategic indicators are the business results the capability was funded to produce in the first place: the ones named back in the initial capability plan or business case, not a new metric invented after the fact because it looks better.

A dashboard that only shows the first tier looks rigorous and proves nothing. A plan that waits for the third tier alone won't surface a problem for a year. The middle tier is what connects the other two, and it's worth naming explicitly instead of assuming it will show up if the other two look healthy.

Even with all three tiers in place, one hard problem remains.

The Attribution Problem

Strategic indicators tend to move for more than one reason at once. They can be affected by market conditions, other initiatives running in parallel, seasonal effects, a competitor's misstep. A capability program can look successful because the metric moved, or unsuccessful because it didn't, for reasons that have nothing to do with the capability at all.

Perfect attribution isn't realistic outside a controlled experiment, and holding out for it is usually just an excuse to avoid measuring. What is realistic: staggering rollout across teams or regions so there's a comparison group, watching whether the leading and capability indicators move before the strategic indicator does rather than after, and being explicit that the case being made is directional and evidence-based rather than a laboratory result. A capability team that says "this is our best read on causation, and here's the evidence" earns more credibility than one that claims certainty it can't back up.

None of this evidence is useful, though, if no one decided in advance what it needed to show.

Define Proof Before You Build, Not After

The habit of defining success at the start, rather than backfilling a definition once results are in, applies with even more force to the measurement plan itself. Before a capability initiative launches, it's worth naming which strategic indicator it's ultimately accountable to, what the leading and capability indicators will be, and what result, in what timeframe, would count as evidence the thing worked.

Deciding this after the effort is already underway invites a familiar failure: the metrics that get chosen are the ones the program is already doing well on. Deciding beforehand keeps the measurement plan honest, and it gives the capability owner a defensible answer the next time someone asks whether the investment paid off.

That discipline holds regardless of what tools are involved in the initiative. But AI changes what's possible in how that evidence gets gathered and read.

Where AI Changes the Validation Game

AI is genuinely useful here, in the same way it's useful earlier in the capability lifecycle. However, it also carries a matching risk that's worth naming directly.

On the useful side, AI can synthesize the three tiers of evidence continuously rather than at quarterly checkpoints, surfacing a capability that's quietly eroding (e.g., usage dropping, quality slipping, rework climbing), months before a lagging business metric would show it. It can also help control some of the noise in the attribution problem, correlating leading indicators against strategic outcomes across teams and time in a way that's genuinely hard to do by hand.

The risk is measurement theater: a real-time dashboard that feels more rigorous than a quarterly spreadsheet but is still just tracking enrollment and usage dressed up in better visuals. A metric that people know is being watched tends to get optimized rather than genuinely improved, and AI-generated dashboards carry an air of objectivity that can make a bad metric harder to question, not easier. The fix isn't avoiding AI in measurement. It's the same fix from the capability strategy itself: know what you're trying to prove before you build the instrument that measures it.

One more piece keeps this from quietly falling apart over time.

Who Owns the Scorecard

Someone must own the measurement plan, and it works best when that person isn't the same person whose funding or reputation depends on the numbers looking good. That's not a statement about anyone's integrity. It's just a structural conflict of interest that's easy to design around and costly to ignore.

Capabilities also decay. That's the same erosion we flagged on the ownership side in Building Capabilities on Purpose, and it's why validation isn't a one-time event at the end of a rollout. Revisiting the chain of evidence on a regular cadence, not just at launch, is what catches the quiet erosion between big checkpoints, before it shows up as a strategic result that mysteriously stopped moving.  

The Bottom Line

The organizations that get this right aren't the ones with the most sophisticated dashboards. They're the ones that decided, before the capability was built, exactly what evidence would prove it worked and exactly which strategic result it was accountable to. Everyone else ends up with a program that was busy, well-attended, and ultimately unable to answer the one question that mattered.

Strategy first. Capability second. Proof running underneath both, the whole way through. If you’re trying to figure out whether a capability investment actually moved the needle, or how to put that proof in place before the next one launches, we’d welcome the conversation.  

Percipio Icon
Challenges
Services
No items found.
Solutions

Most organizations that get the sequencing right, strategy first, then capability, then AI where it earns its place, eventually run into a different problem. Someone in leadership asks a simple question: did it work? And the honest answer, more often than anyone would like, is "we're not sure."

That's not for lack of data. Capabilities built on a solid strategic foundation still generate enrollment numbers, completion rates, satisfaction scores, and usage dashboards for the tools that came with them. The problem is that none of those numbers answer the question being asked. They describe activity. They don't say whether the organization can now reliably do the thing the strategy required, or whether that capability moved the business outcome it was funded to produce.

Getting from "we built the capability" to "it worked" requires a different kind of discipline than building the capability in the first place. It's worth being just as precise about validation as about the initial capability definition.

Two Different Questions, Not One

"Did it work" is actually two questions wearing one sentence, and most measurement efforts quietly answer only the easier of the two.

The first question is whether the capability itself got built: can the organization reliably produce the outcome without depending on the two or three people who happened to be strongest in the room, and does the process, tooling, and incentive structure around the work actually support it. The second question is whether that capability, now built, moved the strategic result it was meant to serve: the new market entered, the operating model that got faster, the regulatory exposure that got smaller.

The initiative to develop a capability can often pass the first test and fail the second. Teams may become genuinely capable at something the strategy or business no longer needs, especially if the initiative ran long enough for the business context to shift underneath it. Treating these two questions as one question is how a report full of green checkmarks ends up next to a strategic result that never moved. Picture a team that spends a year building a sharp analytics capability. They are well-staffed, well-documented and genuinely reliable, but it’s for a market that the business has already decided to exit. The first test passes cleanly. The second one fails just as cleanly, and no dashboard built on leading indicators alone would catch it.  

With the two questions separated, the next challenge is finding measures that can actually answer each one.

Build a Chain of Evidence, not a Single Metric

No single number carries enough weight to prove a capability initiative worked. What holds up under scrutiny is a chain of evidence across three tiers, each one a necessary check on the others.

  • Leading indicators show whether the new behavior is actually happening, for example, is the new process being followed, is the tool being opened, is the decision being made the new way instead of the old way. These often show up more quickly, but they are also the easiest to game. They become vanity metrics. They can be useful as an early warning system but aren’t sufficient as proof of anything on their own.
  • Capability indicators show whether the organization can produce the outcome reliably and without relying on specific individuals: consistency of output quality across people and teams, how the work holds up when a key person is out or moves on, and how much rework or escalation the new way of working still generates. This is the tier most measurement plans skip, and it's the one that distinguishes organizational capability from individual skill.
  • Strategic indicators are the business results the capability was funded to produce in the first place: the ones named back in the initial capability plan or business case, not a new metric invented after the fact because it looks better.

A dashboard that only shows the first tier looks rigorous and proves nothing. A plan that waits for the third tier alone won't surface a problem for a year. The middle tier is what connects the other two, and it's worth naming explicitly instead of assuming it will show up if the other two look healthy.

Even with all three tiers in place, one hard problem remains.

The Attribution Problem

Strategic indicators tend to move for more than one reason at once. They can be affected by market conditions, other initiatives running in parallel, seasonal effects, a competitor's misstep. A capability program can look successful because the metric moved, or unsuccessful because it didn't, for reasons that have nothing to do with the capability at all.

Perfect attribution isn't realistic outside a controlled experiment, and holding out for it is usually just an excuse to avoid measuring. What is realistic: staggering rollout across teams or regions so there's a comparison group, watching whether the leading and capability indicators move before the strategic indicator does rather than after, and being explicit that the case being made is directional and evidence-based rather than a laboratory result. A capability team that says "this is our best read on causation, and here's the evidence" earns more credibility than one that claims certainty it can't back up.

None of this evidence is useful, though, if no one decided in advance what it needed to show.

Define Proof Before You Build, Not After

The habit of defining success at the start, rather than backfilling a definition once results are in, applies with even more force to the measurement plan itself. Before a capability initiative launches, it's worth naming which strategic indicator it's ultimately accountable to, what the leading and capability indicators will be, and what result, in what timeframe, would count as evidence the thing worked.

Deciding this after the effort is already underway invites a familiar failure: the metrics that get chosen are the ones the program is already doing well on. Deciding beforehand keeps the measurement plan honest, and it gives the capability owner a defensible answer the next time someone asks whether the investment paid off.

That discipline holds regardless of what tools are involved in the initiative. But AI changes what's possible in how that evidence gets gathered and read.

Where AI Changes the Validation Game

AI is genuinely useful here, in the same way it's useful earlier in the capability lifecycle. However, it also carries a matching risk that's worth naming directly.

On the useful side, AI can synthesize the three tiers of evidence continuously rather than at quarterly checkpoints, surfacing a capability that's quietly eroding (e.g., usage dropping, quality slipping, rework climbing), months before a lagging business metric would show it. It can also help control some of the noise in the attribution problem, correlating leading indicators against strategic outcomes across teams and time in a way that's genuinely hard to do by hand.

The risk is measurement theater: a real-time dashboard that feels more rigorous than a quarterly spreadsheet but is still just tracking enrollment and usage dressed up in better visuals. A metric that people know is being watched tends to get optimized rather than genuinely improved, and AI-generated dashboards carry an air of objectivity that can make a bad metric harder to question, not easier. The fix isn't avoiding AI in measurement. It's the same fix from the capability strategy itself: know what you're trying to prove before you build the instrument that measures it.

One more piece keeps this from quietly falling apart over time.

Who Owns the Scorecard

Someone must own the measurement plan, and it works best when that person isn't the same person whose funding or reputation depends on the numbers looking good. That's not a statement about anyone's integrity. It's just a structural conflict of interest that's easy to design around and costly to ignore.

Capabilities also decay. That's the same erosion we flagged on the ownership side in Building Capabilities on Purpose, and it's why validation isn't a one-time event at the end of a rollout. Revisiting the chain of evidence on a regular cadence, not just at launch, is what catches the quiet erosion between big checkpoints, before it shows up as a strategic result that mysteriously stopped moving.  

The Bottom Line

The organizations that get this right aren't the ones with the most sophisticated dashboards. They're the ones that decided, before the capability was built, exactly what evidence would prove it worked and exactly which strategic result it was accountable to. Everyone else ends up with a program that was busy, well-attended, and ultimately unable to answer the one question that mattered.

Strategy first. Capability second. Proof running underneath both, the whole way through. If you’re trying to figure out whether a capability investment actually moved the needle, or how to put that proof in place before the next one launches, we’d welcome the conversation.  

Ready to work with us?