Beyond Engagement: Measuring Whether AI Actually Develops Human Capability
Why DAU, retention, and session length are the wrong metrics for development products — and what to measure instead.
Deep Dive · Next AI Labs · Human-Agent Research · 14 min read
Most AI products measure engagement. Daily active users, session length, retention curves — the standard toolkit borrowed from consumer software and applied wholesale to products that claim to develop human capability.
The problem is that engagement metrics cannot distinguish between a user who is developing capability and a user who is developing dependency. Both look identical in the dashboard: frequent sessions, long interactions, consistent return visits. One is growing. The other is outsourcing thinking they used to do themselves. The metric says "success" in both cases.
This is the central tension for anyone building AI that claims to develop human capability. The metrics that feel good and the metrics that mean something are not just different — they are often inversely correlated. And the entire industry is optimizing for the wrong ones.
---
The ambiguity runs in both directions. A user who opens the app daily and generates hundreds of coaching conversations could be deeply engaged in genuine development work — or could be outsourcing their own thinking to the system. A user whose session frequency is declining could be churning — or could be internalizing what they learned and needing less scaffolding, which is the actual goal of a development product.
Engagement metrics cannot distinguish between these cases. DAU, session length, and retention measure one thing: whether the user came back. They say nothing about whether the user grew. This is not an edge case. It is the default failure mode when engagement vs development is not explicitly addressed in the measurement system. The user who gradually needs less support has declining engagement metrics and maximum capability development. The user who asks the AI to solve their problems every day has great engagement metrics and zero capability development.
The problem compounds. When a team optimizes for engagement, the product learns to maximize return visits. In a coaching context, this means the system learns to be helpful enough to bring you back but never transformative enough to make itself unnecessary. This is goodharts coaching trap — the coaching equivalent of a tutor who keeps you dependent so you keep booking sessions.
---
You cannot directly measure "did this human become more capable?" It is not a number you can query from a database. Capability development is longitudinal, domain-specific, deeply contextual, and mostly invisible to the systems that are supposed to produce it.
So you need proxies — measurable signals that stand in for the thing you actually care about. The question is which proxies, and how do you prevent goodharts coaching trap from corrupting them.
A capability proxy is only useful if it has three properties. First, it must actually correlate with the thing it represents. Session count does not correlate with development. Specificity of coaching language does. Second, it must be resistant to gaming — both by the system (which will optimize whatever you measure) and by the team (which will celebrate whatever turns green). Third, it must degrade gracefully. When the proxy stops measuring what you care about, you need to notice before you have spent a quarter optimizing in the wrong direction.
This is harder than it sounds. Most teams skip the proxy analysis entirely and measure what is easy: sessions, clicks, time-on-page. These are not wrong metrics. They are metrics for a different question. They answer "are people using this?" when the real question is "is this working?"
---
We arrived at our current measurement system not through a single design process but through a sequence of failures. Each time we optimized for a metric and watched it diverge from what we actually cared about, we added a new layer. The result is imperfect and still evolving, but it reflects something closer to "is this working?" than anything we started with.
The core measurement instrument evaluates coaching quality across 10 dimensions, each scored 1-10 with specific behavioral anchors. This is not engagement measurement — it is output quality measurement.
The 10 criteria: emotional intelligence, context depth, re-entry bridge, timing, voice consistency, subject line craft, anti-spam compliance, model quality, template polish, and feedback loop integration.
Two of these reveal the philosophy most clearly.
Emotional Intelligence is designed to punish exactly the behavior that engagement-optimized systems produce. Each criterion has behavioral anchors across the 1-10 scale. The spectrum runs from generic coaching language at the low end — output that could be sent to any user — through to coaching that uses the user's own words, references their specific tensions, and connects dots the user did not explicitly state. The highest scores reflect intuition calibration: the system has developed enough of a working model of the user to surface things they have not said directly.
Context Depth measures something subtler — not what the system knows, but what it does with what it knows. Low scores indicate database retrieval formatted as coaching: listing goals and session counts with no temporal awareness. High scores indicate genuine synthesis — weaving multiple dimensions, noticing patterns the user did not explicitly state, referencing what they have not done alongside what they have. This is the signal that the system is modeling the user's development trajectory, not just their activity log.
The gap between the middle and top of these scales is the gap between a system that has your data and a system that has developed a working model of your development. Both produce coaching output. One produces development.
The calibration framework has a companion: an explicit ban list. These phrases score zero on emotional intelligence and trigger automatic flagging:
"I get it." "You've got this." "Keep up the great work." "Cheering you on." "Rooting for you."
These are the hallmark of a system pretending to care. They are engagement-optimized language — they feel warm, they fill space, they validate without challenging, and they produce zero development. An AI coaching system that defaults to "You've got this" is the equivalent of a personal trainer who just counts your reps and says "nice job" — technically present, functionally useless.
This is a philosophical stance embedded in a measurement system. Genuine development requires honest reflection, not cheerful validation. If the system cannot find something specific and real to say, it should say less, not fill the gap with encouragement theater.
Separate from the output calibration, we score the quality of the context available to the system before it generates anything. The health scoring rubric rates 1-10 with weighted signals:
| Signal | Weight | What It Measures | |--------|--------|------------------| | Sessions | 0.25 | Depth of coaching history | | Goals | 0.25 | Clarity of development direction | | Recency | 0.20 | How current the relationship is | | Memories | 0.15 | Accumulated understanding over time | | Funnel stage | 0.10 | Where they are in their journey | | Intake | 0.05 | Initial self-assessment data |
This is not engagement measurement. It is relationship depth measurement. A user with high session count but no goals and no memories scores lower than a user with fewer sessions but clear goals and rich memory accumulation. The scoring answers a question engagement metrics cannot: how much does the system actually know about this person's development journey?
The practical consequence: when context quality is low, the system's output quality ceiling drops regardless of how sophisticated the model is. You cannot produce a 9 on emotional intelligence with a 3 on context depth. This creates a natural pressure to build systems that accumulate genuine understanding rather than just interaction count.
Here is where capability proxy meets business reality.
Across 5,000+ users and 30,000+ coaching sessions, the system produces a 28x ratio of customer acquisition cost to lifetime value. This is not presented as a business metric. It is presented as evidence.
If the system were producing engagement without development — if users were returning out of habit or dependency rather than genuine growth — the ratio would not hold. Engagement-addicted users churn when the novelty fades. Dependent users churn when they realize they are not growing. Only users who are genuinely developing capability sustain the kind of long-term value relationship that produces a 28x multiple.
The 28x is a capability proxy by inversion: instead of measuring development directly (which we cannot do), we observe its financial shadow. Humans who find genuine value in their development — who experience real shifts in how they approach problems, relationships, and decisions — stay and pay at rates that are anomalous for consumer software. The anomaly is the evidence.
This is not a perfect proxy. It conflates development with perceived development, and perceived development with willingness to pay. These are not identical. But it is a better proxy than DAU, which conflates presence with growth. At least the 28x requires sustained value perception over time rather than just repeated opening of an app.
The proactive coaching re-engagement system — which sends personalized outreach to users who have drifted away — produces a 7:1 revenue gain-to-loss ratio when sustained engagement is maintained over time.
The interesting finding is not the ratio itself but its mechanism. The re-engagement system does not retain users through urgency, discounts, or FOMO — the standard retention playbook. It re-connects them to development work they actually value. It references their specific goals, their progress trajectory, what they were working on when they disengaged.
The system also surfaces an important dynamic: users who were subscribing out of inertia — zombie subscribers generating revenue but receiving no value — tend to cancel when the coaching outreach makes them confront whether they are actually using the system for growth. This is a baseline drift correction: the re-engagement emails force a recalibration between "I should be doing this" and "I am actually doing this." Proactive coaching outreach brings genuine users back while low-value subscribers either re-engage meaningfully or cancel.
A pure engagement optimization would try to prevent these cancellations. We treat them as signal. If a user cancels because an honest coaching email made them realize they were not actually developing, the system is working correctly. The 7:1 ratio is high precisely because the remaining users are genuine — they did not just forget to cancel.
Underneath all of these measurements is a strategic principle: the search for transformational fulcrums — small interventions that unlock outsized shifts in user outcomes.
Every product initiative is evaluated not by engagement impact but by leverage at these fulcrums. Does improved emotional attunement in coaching output change how a user approaches a specific real-world problem? Does better sensemaking — helping a user see patterns in their own behavior they had not noticed — produce a shift that persists after the session ends?
This is the deepest layer of engagement vs development. Engagement asks: did they come back? Development asks: did something change in how they operate? The fulcrum framing forces every feature to articulate a theory of change — not "users will spend more time" but "users will notice X about themselves, which will shift Y in how they approach Z."
We cannot yet prove most of these theories. But having them is qualitatively different from not having them. A team that ships features with engagement theories ("this will increase session length") builds different things than a team that ships features with development theories ("this will help users notice when they are avoiding a hard conversation").
---
This section exists not as performative humility but as genuine cartography of what we do not know. These are real open questions, not softballs.
We do not yet measure whether users are actually developing capability over time in any rigorous sense. We measure whether the system's output is high quality (the calibration framework) and whether users stay (the 28x LTV). But the link between "the system said something insightful" and "the human changed how they approach problems" is unobserved.
What we want: time series measurement of specific capability dimensions. Does a user who works on leadership communication through coaching actually communicate differently six months later? Does a user who explores emotional regulation in coaching sessions show measurable changes in how they process difficult situations?
What blocks us: measurement requires access to the user's life outside the product. Self-report is unreliable. Behavioral proxies within the product (like changes in how they describe their challenges) are suggestive but not definitive. This is the measurement aspiration — the measurement we would build if we could.
The most important question in coaching — did this conversation change anything in the real world? — is the one we are least equipped to answer. A user could have a profound insight during a coaching session and do absolutely nothing with it. A different user could have a seemingly ordinary session and make a life-altering decision that afternoon. We see the sessions. We do not see the aftermath.
Some users voluntarily report transfer: "I tried that thing we discussed and it worked." These reports are gold, but they are self-selected and sparse. Building systematic measurement of transfer — without invasive surveillance of users' actual lives — remains unsolved.
The outcome we ultimately care about — is this person flourishing more because of their engagement with the system? — is measurable only through self-report. And self-report in the context of a product the user is paying for is contaminated by confirmation bias, sunk cost rationalization, and desire to believe the investment is working.
We have considered periodic flourishing surveys based on validated instruments from positive psychology. We have not implemented them because we are not confident we can interpret the results without the confounds overwhelming the signal. A user who reports higher well-being after using coaching for six months may be experiencing genuine growth, placebo effect, or simply the passage of time. Distinguishing these requires controlled study designs we do not have the scale or resources to run.
---
The 28x LTV is not despite the philosophical stance against engagement metrics. It is because of it.
This is the claim that most product teams find counterintuitive, so it is worth being explicit about the mechanism.
When you optimize for engagement, you build features that bring people back. When you optimize for development, you build features that genuinely help people. These are different features. The engagement-optimized feature set includes notifications, streaks, gamification, social pressure, and content that is interesting but not challenging. The development-optimized feature set includes honest assessment, uncomfortable questions, personalized challenge calibration, and output that is sometimes hard to hear.
The engagement feature set produces higher short-term metrics and lower long-term value. Users engage more initially but the relationship is shallow — they are interacting with the product's retention mechanics, not its core value. When a competitor offers slightly better retention mechanics, they leave.
The development feature set produces lower short-term metrics and higher long-term value. Users may engage less frequently — because they are doing real work between sessions, not just consuming — but their relationship with the product is deep. They are not retained by mechanics. They are retained by genuine value. They are the users who refer others, who write unsolicited testimonials, who stay through price increases and feature changes.
The 28x ratio is the financial expression of this difference. It is what happens when a critical mass of users are staying because the product actually works, not because the product is good at making them come back.
The 7:1 ratio reinforces the point from the other direction. The re-engagement system works not by manipulating users into returning but by honestly re-connecting them to development work they care about. The users who come back are coming back to something real. The users who leave are honestly assessing that they are not using the system for growth. Both outcomes serve the goal better than a retention campaign that brings everyone back and measures the resulting DAU spike as success.
---
Three hard questions we are actively wrestling with. Not theoretical — these affect current product decisions.
1. Is declining engagement actually a signal of success? (Hypothesis, not established finding)
We hypothesize that a user who gradually needs less coaching is developing capability, while a user who increases usage may be developing dependency. But we have not rigorously tested this claim. It is possible that our most engaged users are also our most-developing users — that engagement and development are correlated, not anti-correlated, and we have constructed a framework that makes a virtue of churn. Testing this requires measuring development independently of engagement, which brings us back to the gap we described above. Until we can measure development directly, the relationship between engagement and development is a hypothesis, not a finding.
2. Are our capability proxy signals measuring development or measuring selection? (Open question)
The 28x LTV and the 7:1 ratio may not reflect development at all. They may reflect the fact that our product attracts — and retains — a specific personality type: self-directed, growth-oriented, willing to pay for self-improvement. These users would produce high LTV in almost any personal development product, regardless of whether it actually develops them. If our "development signal" is actually a selection effect — the product filtering for a certain type of user rather than developing them — then our metrics are measuring who shows up, not what happens to them. We do not currently have the data to distinguish between these explanations.
3. Can the calibration framework be gamed by a sufficiently sophisticated model?
The 10-criterion framework scores output quality on dimensions like emotional intelligence and context depth. But these are assessed based on surface features of the output — does it use the user's own words, does it reference patterns the user did not state, does it avoid banned phrases. A model that learns to mimic these surface features without genuine understanding would score well on the framework while producing coaching that is sophisticated-sounding but ultimately hollow. As models improve, the gap between "sounds like a 9" and "is a 9" may widen. We have no current mechanism for detecting this divergence. The normalized signal approach helps with consistency, but not with the fundamental question of whether surface quality reflects actual coaching depth.
---
The honest summary: we have built a measurement system that is better than engagement metrics, but we have not built a measurement system that directly measures what we care about. The calibration framework measures output quality. The context quality scorer measures relationship depth. The LTV ratio measures sustained value perception. The re-engagement ratio measures genuine re-connection. None of these directly measure human development.
What we have done is construct a capability proxy stack — multiple signals, each imperfect, each measuring a different facet of the underlying phenomenon. When they all point in the same direction (high calibration scores, deep context, sustained retention, honest re-engagement), the case for genuine development is stronger than any single metric would provide. When they diverge — high retention with low calibration, or strong re-engagement with shallow context — that divergence is itself a signal worth investigating.
The measurement aspiration remains: to measure not whether people use the system, and not even whether the system's output is good, but whether humans are genuinely developing capability through their interaction with it. This requires solving the transfer problem, the longitudinal measurement problem, and the self-report contamination problem. These are research-grade challenges, not engineering challenges. We do not pretend to have solved them.
What we do claim is that the measurement philosophy matters — that a team which asks "is this working?" will build different things than a team which asks "are they using it?" — and that the difference compounds over years of product decisions into fundamentally different products. One produces engagement. The other produces development. And the business outcomes suggest that the market, eventually, knows the difference.
Practice this with IX Coach
IX Coach brings these alignment principles into a guided, adaptive coaching experience.
7 days free, then $40/month (~$1.30/day).