top of page

The Variable Most Leadership Programs Never Measure

Why capable executives revert after development programs — and what the Rawe Adaptive Leadership Framework does differently


A pattern every senior HR leader recognizes


The program was well designed. The facilitators were credible. The evaluation scores were strong. And six months later, the executives who completed it are leading almost exactly as they did before. The feedback conversations are still avoided, the delegation still collapses under pressure, and the strategic patience discussed so thoughtfully in the cohort room has not survived contact with the quarter.


This is not an isolated failure. Leadership development is a global industry estimated at more than $90 billion, and multi-source studies consistently report that between half and three-quarters of programs fail to produce sustained behavioral change. The conventional response treats this as an execution problem: better content, more engaging delivery, longer programs, stronger follow-up. Organizations have been running that experiment for decades. The failure rate has not moved.


The evidence points to a different explanation. The problem is not how leadership development is delivered. It is what leadership development measures — and, more precisely, what it fails to measure before it begins.


The mechanism: skills need a structure to hold them


Decades of research in adult development — most notably Robert Kegan’s constructive-developmental theory — established something the leadership development industry has largely ignored: adults differ not only in what they know, but in the underlying structure through which they make meaning of authority, feedback, conflict, and identity. That structure develops throughout adulthood and determines what a leader can reliably do under pressure, regardless of what they have been trained to do.


Consider two executives who complete the same program on leading through disagreement. Both can articulate the model. Both perform well in the practice sessions. Back at work, one of them applies it consistently — including when the disagreement comes from a superior whose approval matters. The other applies it selectively, and abandons it entirely the moment the conflict threatens an important relationship. The difference is not motivation, intelligence, or effort. For one leader, the approval of key stakeholders is something they have — a factor they can weigh. For the other, it is something they are — a structure they cannot yet step back from and examine. No amount of skills training changes that by itself.


This is the mechanism at the center of the Rawe Adaptive Leadership Framework (RALF): a leader’s developmental structure functions as a moderating variable that determines whether competency-based development produces durable behavioral change. Skills are the contents. Meaning-making structure is the container. Most programs pour new content into containers they have never assessed — and then attribute the spillage to poor engagement or weak reinforcement.


What is RALF?


RALF is a research-grounded framework for executive and organizational development. It is an applied synthesis — it integrates constructive-developmental theory, complexity leadership theory, adaptive leadership scholarship, and reflective practice research with original empirical findings from a 2024 doctoral study of human resource development professionals. It does not claim to be a new theory of adult development. It claims to organize what the research already supports into a system an organization can actually run.


The framework rests on three interdependent dimensions of leadership development:


  • Vertical capacity. The developmental structure through which a leader makes meaning — how they relate to authority, feedback, competing perspectives, and their own assumptions. This is the dimension most programs never assess.


  • Horizontal capability. The skills, knowledge, and behaviors the role requires — the familiar territory of competency models, training, and 360° feedback. Necessary, but effective only when the underlying structure can support it.


  • Reflexive practice. The active capacity to examine one’s own meaning-making in real time — the bridge between the other two dimensions and the mechanism through which structural growth actually occurs.


Development that moves in only one dimension is fragile. A leader with strong skills and an unexamined structure reverts under pressure. A leader doing deep reflective work without capability development has insight they cannot operationalize. RALF is built to move all three together.


Growth in these dimensions becomes visible through five adaptive capacities identified in the underlying research as decisive in post-pandemic leadership environments: adaptive customization, reflexive self-awareness, multi-generational responsiveness, holistic development orientation, and a servant orientation to influence. These are observable behaviors, not personality traits — which means progress against them can be seen, described, and tested rather than merely asserted.


Why some leaders grow in a program and others plateau


Watch a cohort closely, and a contrast emerges. Some participants treat a challenge to their thinking as information: they can hold their own framework at arm’s length, examine it, and revise it. Others experience the same challenge as a threat to be managed — politely deflected, privately dismissed. Some can genuinely coordinate competing stakeholder perspectives without losing their own position; others can only serially adopt whichever perspective is in the room. These differences are structural, assessable, and predictive of which participants a given program design can actually serve.


The practical implication for an HR leader is uncomfortable but clarifying: the same program is not the same intervention for every participant. A design calibrated to one developmental position over-challenges some leaders into shutdown and under-challenges others into boredom. This is why aggregate program evaluations look fine while behavioral outcomes disappoint — the average hides a calibration failure that was baked in before the first session.


Bright open-plan office with several coworkers at desks, talking and working amid brick walls, bookshelves, and sunlight

How the framework runs: Phase 0 and the developmental loop


RALF translates assessment into development through a recursive four-phase loop — Discover, Disrupt, Design, Demonstrate — that governs both individual executive coaching and cohort programs. Each completed cycle begins from a more developed platform; the loop is a practice architecture, not a one-time event.


The loop is entered through Phase 0: a structured organizational diagnostic conducted before any development activity begins. Phase 0 assesses the developmental demands of the roles in question, the leaders' current developmental positions within them, and the gap between the two. It is the variable most leadership development programs omit — and its omission is why so many well-executed programs fail. Without it, program design is calibrated to an assumed rather than an assessed developmental position. Phase 0 is a standalone, purchasable engagement in its own right: organizations frequently discover that the diagnostic alone reframes a problem they had been treating as an engagement, retention, or skills issue.


Measurement with honest limits


Vertical capacity is assessed through the Leadership Complexity Inventory (LCI™), a proprietary three-part instrument designed to estimate a leader’s developmental zone and the gap between that zone and the complexity demands of their role. The LCI™ combines structured self-report with two qualitative methods drawn from established developmental assessment traditions, and it reports a developmental zone estimate — not a score, not a label, and not a verdict.


Two things distinguish this from the assessment landscape most HR leaders know. First, the LCI™ measures developmental structure, not personality type or behavioral style — it asks a different question than the instruments already in your toolkit, and it is designed to complement rather than replace them. Second, the instrument is in active psychometric validation, and we say so plainly. The validation roadmap — expert content review, convergent validity testing against the established Subject-Object Interview, inter-rater reliability studies — is published, sequenced, and underway. In a field crowded with instruments that promise everything and disclose nothing, a falsifiable claim with a public validation agenda is not a weakness. It is the point. Any framework that cannot tell you how it could be proven wrong cannot tell you why it should be believed.


What this changes for your organization


RALF does not promise transformation. It changes the quality of the decisions you make about development — and decision quality is where the waste in this category actually lives. Concretely, a developmental approach gives an executive team and its HR leadership the ability to:


Sequence investment correctly — identifying which leaders are positioned to convert a given program into durable change now, which need a different developmental intervention first, and which are already operating beyond what the planned program would offer.


Read succession and stretch-role risk structurally — assessing whether a candidate’s developmental position matches the complexity demands of the target role, rather than inferring readiness from performance in a less demanding one.


Diagnose before prescribing — treating turnover, engagement collapse, and failed change initiatives as potential symptoms of a developmental mismatch between leaders and role demands, rather than reflexively purchasing another program.


Hold vendors to a mechanism — asking of any proposed development investment the question this framework is built on: through what causal pathway is this intervention expected to produce durable change, and for whom?


Where to start


If your organization has run capable programs that produced strong evaluations and weak behavioral change, the explanation is unlikely to be effort, content, or facilitation. It is far more likely to be an unmeasured variable — one that has been sitting underneath every development investment you have made.


Phase 0 exists to measure it. It is a bounded, standalone diagnostic engagement, and everything else in the RALF system — executive coaching, cohort development, organizational consulting — is calibrated by what it finds. Begin with the diagnostic. The rest of the work depends on it.


Comments


bottom of page