Adding metrics does not necessarily improve understanding. Many dashboards contain several measures of the same underlying behaviour. Metrics become more diagnostic when combined as mutually destructive pairs: cost with quality, containment with sentiment, or accuracy with latency.
The right pairs change as an AI product moves from pilot to production. Measurement should follow the decisions the team needs to make at each stage.
Operational monitoring and strategic measurement serve different purposes. Teams may track hundreds of system signals while using only a few pairs to guide product decisions.
Headline metrics can tell a persuasive story while leaving an important product decision unresolved.
Klarna’s AI assistant was publicly associated with greater throughput, lower costs, and customer-satisfaction scores comparable with human agents. The following year, the company decided to increase access to human support, with its chief executive acknowledging that cost reduction had received too much emphasis.
This was not a rejection of the AI assistant or its underlying technology. It was an adjustment to the balance between automation and human service as the company learned from operating the product. As automation took on more of the high-volume, simpler queries, the human agents Klarna needed were those equipped for complex, sensitive cases.
Having defined what a product should achieve from the outset, most organisations building AI products will eventually face the same question: is it actually working?
If the answer is unclear, the instinct is often to add more metrics. Three become 10, then 10 become 30. The dashboard grows richer, but the team’s understanding may not improve.
The problem is not always the quality of the individual measures but the relationship between them. Customer satisfaction, Net Promoter Score, reviews, and thumbs-up rates can all provide useful signals, but they may reflect similar changes in overall sentiment. When they move together, they confirm that something has happened without necessarily explaining why.
AI teams should therefore look beyond metrics that agree with each other, and identify measures that expose competing outcomes. These mutually destructive pairs reveal and help to monitor the impact of trade-offs behind product performance allowing for better decision-making.
In one pre-launch deployment, the joint team was developing a real-time AI voice agent for inbound customer-support calls. One of the hardest questions was not about model selection or orchestration. It was how the organisation would know whether the product was working once customers began using it at scale.
The initial framework used three measures:
Containment rate: how often the AI resolves a call without transferring it to a person.
Escalation rate: how often a call transfers to a human agent.
Fulfilment rate: how often the customer’s issue is ultimately resolved.
Each measure was reasonable. Together, however, they could not answer an obvious question: if escalation increases, what does that tell us?
The team decomposed escalation into eight subtypes. It then added abandonment, journey and timing measures, language-understanding scores, and fulfilment by query type. The framework eventually contained 31 metrics across six categories.
It could describe escalation in detail, but it still could not reliably diagnose its cause. Most of the metrics were variations of the same behaviour, so they moved together rather than testing competing explanations.
The dashboard had become observational rather than diagnostic.
What the team needed was not another layer of decomposition. It needed metrics that constrained each other.
We call these mutually destructive pairs: two measures where improving one in isolation can damage the outcome represented by the other. The name describes the failure mode created by one-sided optimisation, not the desired state.
When both sides remain healthy, the product may be operating sustainably. When they diverge, the direction of that divergence helps the team decide where to investigate.
What we had | Mutually destructive pair | What the pair may reveal |
|---|---|---|
Escalation rate divided into eight subtypes | Escalation rate ↔ time to escalation | Immediate escalation may indicate a trust or framing problem; later escalation may indicate that the system cannot complete the task. |
Containment rate and fulfilment rate reported separately | Containment rate ↔ customer sentiment | Whether containment represents a satisfactory resolution or a customer abandoning the attempt. |
Fulfilment rate by intent type | Fulfilment rate ↔ conversation depth | Whether successful resolution is efficient or requires an exhausting interaction. |
Looking deeper into how this is exposed and what to do with this insight, let’s consider escalation rate and time to escalation. The team will not know how customers behave until real calls arrive, but it can define the hypotheses it needs to test.
If more calls start being escalated and customers leave the AI experience within the first 30 seconds, the team should investigate trust, disclosure, tone, and opening interactions. If customers escalate after spending several minutes attempting a task, capability or workflow coverage becomes the more likely problem.
The headline escalation number is the same. The product decision is different.
A useful pair does not prove the cause on its own. It narrows the investigation and makes the next decision clearer.
The Klarna example introduced earlier shows how this principle applies when cost and service quality interact. It demonstrates how an AI-enabled operating model can evolve as a company monitors the impact of trade-offs and learns from its deployment.
In February 2024, the company reported that its AI assistant had handled 2.3 million conversations in its first month, performed work equivalent to 700 full-time agents, and achieved customer-satisfaction scores comparable with human agents. Klarna estimated that the assistant would contribute $40 million in profit improvement during 2024. These were Klarna’s own reported results rather than an independent evaluation.
In May 2025, Klarna’s chief executive said the company had placed too much emphasis on cost reduction in customer service and described plans to increase access to human support. This represented an adjustment to the balance between automated and human service, rather than a rejection of the AI assistant or its underlying technology.
The public evidence illustrates why efficiency measures should be considered alongside the needs of different customers and interactions. An AI system may perform well on average while some complex, sensitive, or unusual cases still benefit from an accessible human channel.
Monitoring both sides of that relationship helps a company decide where automation creates value, where human support remains important, and how the balance should change as new evidence emerges.
Other mutually destructive pairs across AI products may include:
Mutually destructive pair | Risk it helps expose |
|---|---|
Response accuracy ↔ response latency | A system that is technically accurate but too slow for the workflow. |
Task completion ↔ user override rate | An AI workflow that completes tasks which users repeatedly redo. |
Cost per interaction ↔ evaluated output quality | Savings achieved by degrading the customer or employee experience. |
Adoption ↔ time to value | Growth in sign-ups without corresponding user value. |
The purpose is not to make both measures increase indefinitely. It is to make the trade-off visible before one-sided optimisation creates an operational problem.
A related challenge appeared in a player-support deployment for a mobile-games company. The system handled high-volume issues such as lost progress, payment disputes, and account access.
Efficiency measures mattered because the system operated at scale. But player support is not simply an operational queue. Players often arrive frustrated because something has already gone wrong elsewhere in their experience.
That starting point changes how customer-satisfaction data should be interpreted. A player whose issue is resolved correctly may still report low satisfaction because they lost progress in the first place. Reading that score without context can penalise the support interaction for frustration created earlier in the customer journey.
The team therefore needed to distinguish the customer’s starting sentiment from the effect of the support experience. The more useful question was not, “Was the player happy?” It was, “Did the interaction improve the situation relative to where the player started?”
That comparison can help separate product frustration from support quality, provided the team has a reliable way to measure both.
AI products change, but their metrics often remain fixed.
During a pilot, the central question may be whether the system is trustworthy enough to justify continued investment:
Does it reliably complete the core task?
Do users trust it enough to continue?
How does it behave outside the most common scenarios?
Can failures be identified and recovered safely?
Those questions favour pairs such as:
core-task success ↔ edge-case performance;
automation rate ↔ human override rate; and
completion speed ↔ user confidence.
Once the product becomes operationally important, the questions change:
Can it scale without reducing quality?
Do its economics improve with use?
Does performance remain stable as adoption grows?
Are human interventions occurring in the right places?
The corresponding pairs may shift towards:
cost per interaction ↔ evaluated output quality;
adoption breadth ↔ depth of use; and
automation rate ↔ operational risk exposure.
Early metrics are not necessarily wrong. They answer the questions that mattered at an earlier stage.
The risk comes during the transition. Pilot metrics often persist because teams know how to report them and nobody owns the decision to retire them. Measures that once supported learning can gradually become vanity metrics.
Pairs should therefore have a lifecycle. Teams should introduce them for a defined decision, review whether they still expose a material trade-off, and retire them when the product or decision changes.
AI systems require detailed observability, alerting, quality assurance, and evaluation. Removing those signals would make the product harder to operate safely. But operational monitoring is not the same as leadership measurement.
Monitoring helps teams detect incidents, trace failures, and understand system behaviour. Decision metrics help product and business leaders decide whether to invest, intervene, change direction, or accept a trade-off.
An organisation may monitor hundreds of technical and operational signals while elevating only two or three mutually destructive pairs for a particular product decision. Keeping that decision layer small makes prioritisation easier.
The appropriate review cadence depends on the product. A new or rapidly changing system may require weekly decision reviews, while a mature product may support a monthly or quarterly cadence. The principle is more important than the interval: review the pair often enough to act before the trade-off becomes expensive or unsafe.
A pair only becomes useful when the organisation agrees what will happen if it deteriorates.
That requires more than defining a red line for divergence. Teams should consider three conditions:
Absolute failure: one measure crosses an unacceptable threshold, regardless of the other.
Divergence: one measure improves while its counterbalance deteriorates.
Joint deterioration: both sides decline, suggesting a broader product or operating problem.
Each pair should have:
a named owner;
a clear decision it supports;
agreed thresholds or evaluation criteria;
an investigation path; and
a set of possible interventions.
Without those elements, the organisation is observing the product rather than managing it.
Before adding another measure, choose one important product decision and work through these questions. Write down the answers so the review ends with an agreed next step.
1. Which decision do we need these metrics to help us make?
Be specific: are we deciding whether to expand automation, change a model, or improve human handover? Name the decision before choosing the measures.
2. If this number improves, what could get worse?
Identify the outcome you need to protect and a measure that would reveal harm. For example, pair cost per interaction with evaluated output quality to check whether cheaper answers are still useful.
3. What could the headline numbers be hiding?
Read both measures for the same users, tasks, and time period, then look for groups experiencing worse outcomes. Consider starting conditions too: low satisfaction may reflect frustration that predates the support interaction.
4. What would make us act, and who owns the response?
Set criteria for acting when a measure crosses an unacceptable limit, one improves while the other worsens, or both deteriorate. Agree who will investigate, what they will check first, and when they will report back.
5. Does this pair still fit the product’s current stage?
Decide whether to keep, replace, or retire it. A pilot may focus on task reliability and user confidence; a live service may need closer scrutiny of cost and quality. Set a date to revisit the choice.
AI product measurement should do more than describe performance. It should expose the trade-offs the organisation is making and make the next decision clearer.