Most organizations invest thousands of dollars in leadership training every year, yet studies consistently show that up to 90% of new skills learned in training programs fail to transfer back to the workplace. That is a staggering return on investment problem, and it raises a critical question: what separates leadership development that actually changes behavior from programs that simply check a box?
The research has answers, and they may surprise you. Effective leadership training is not about the length of a program, the prestige of the facilitator, or even the quality of the content alone. It depends on a specific combination of design principles, organizational conditions, and follow-through strategies that most companies overlook entirely.
In this analysis, we break down what the latest research reveals about training transfer, explore the common failure points that undermine even well-designed programs, and outline the evidence-based practices that help leadership skills stick long after the workshop ends. Whether you are designing a new program or evaluating an existing one, this research will give you a sharper lens for making smarter investments in your leaders.
The Money Is Going In. The Results Are Not Coming Out.
Organizations spend an estimated $60 billion annually on leadership development globally. U.S. companies alone spent $102.8 billion on all employee training in 2024 and 2025. Those are not small numbers. They represent boards signing off, budgets expanding, and HR teams building calendars full of workshops, speakers, and online modules. The reasonable expectation is that results follow the spend. They are not following.
Gallup’s State of the Global Workplace data tells the story in three numbers. Global manager engagement stood at 31% in 2022. It fell to 27% by 2024. It then dropped to 22% in 2025, a five-point single-year collapse that Gallup describes as the loss of the engagement premium. That nine-point decline happened while training budgets were growing, not shrinking. Seventy percent of the variance in team engagement stems directly from the manager, which means this is not a morale problem confined to individuals. It cascades to every team those managers lead.
The best-practice data makes the situation harder to excuse. In organisations that develop managers well, 79% of managers are engaged, nearly four times the 22% global average (Gallup, cited by Evalflow). That gap is not structural or inevitable. It is a development and measurement failure. The conditions that produce engaged managers are reproducible. Most organisations are simply not reproducing them.
The skills picture confirms the pattern. Seventy-four percent of business leaders admit they are not keeping pace with the skills their organisations need, despite record training budgets. Content is not the problem. The industry produces more leadership content every year. The problem is structural: training is bought as an event, measured by satisfaction scores on the day, and considered complete. What happens in the ninety days after the training, when the real transfer either occurs or does not, receives almost no attention and almost no investment.
The paradox is specific. Spend is rising. Measurement is absent. Results are declining. That combination does not point to a shortage of courses or facilitators. It points to a fundamental problem in how leadership training is purchased, deployed, and evaluated. Fast Company’s coverage of the Gallup findings frames the manager engagement crisis as an urgent business problem, not an HR statistic. The evidence supports that framing.
Why Training Alone Does Not Transfer to the Job
The research on this is three decades old and unambiguous. Baldwin and Ford’s 1988 review established that training without post-training reinforcement transfers to the job at roughly 5 to 10 percent. Joyce and Showers found that the same content, reinforced with structured coaching after the program, transfers at 80 to 90 percent. That is not a marginal difference. It means the standard delivery model, a two-day workshop followed by nothing, produces approximately one-tenth of the behavioral change that a coached model produces. Organizations have known this since before most of their current HR leaders entered the workforce.
The Evidence Has Not Changed the Default
A peer-reviewed framework published in Behavioural Sciences (October 2024) and indexed on PMC synthesized four separate literature reviews and produced 65 evidence-informed strategies distributed across five program phases: foundation, before the program, during, at conclusion, and after. The concentration of strategies in the pre-program and during phases, combined with a dedicated post-program phase, confirms a consistent finding across the literature: value is made or permanently forfeited based on what happens outside the training room. The same source estimates global annual investment in leadership development at $60 billion. The workplace application of that spend remains typically low.
The industry norm persists regardless. Programs are purchased as events. Success is measured by attendee satisfaction scores collected at the end of day two. No behavioral data is gathered after participants return to their teams. No individual is accountable for whether the learning appeared in observable behavior three months later. This is not a fringe practice; it is the standard purchasing model. 81 percent of leaders believe development should be continuous rather than a one-time event, yet event-based delivery continues to dominate how organizations buy and vendors sell.
The First-Time Manager Gap
The transfer problem has a specific, acute form that most organizations refuse to treat as a structural issue. A person performs well in an individual contributor role. The organization promotes them into people management as a reward for that performance. No preparation follows. No onboarding into the new role, no coaching on how to run a one-on-one, no instruction on how to hold someone accountable or delegate without losing control of outcomes.
The new manager defaults to what they know: individual contribution. They keep doing the work themselves because that is what earned the promotion. Their team waits for direction, feedback, and accountability that does not arrive. The research is clear that 77 percent of organizations report a leadership gap, yet only 48 percent have a formal program to address it. That 29-point gap is not a budget problem. It is a structural failure in how organizations think about the transition from doing work to leading people.
The cost shows up in attrition, in disengaged teams, and in managers who are struggling in private because no one built the post-training conditions that transfer research has required for thirty-seven years.
Satisfaction Scores Are Not a Measurement of Behavior Change
Most organizations measuring leadership training success are measuring the wrong thing. Satisfaction surveys, completion rates, and post-workshop assessment scores tell you whether training occurred. They say nothing about whether anything changed on the job the following Monday, or the Monday after that.
This is not a minor methodological oversight. It is the default practice across a market spending an estimated $60 billion annually on leadership development globally. A manager can sit through a two-day workshop, score well on the closing assessment, give the program a four out of five on the feedback form, and return to managing exactly as before. The satisfaction score looks identical to the score from a program that actually produced behavior change. Without pre- and post-measurement, no one can tell the difference.
The Scrutiny Following the Spend
As leadership development spending accelerates, the people approving the budgets are starting to ask harder questions. CFOs and CEOs are not naturally inclined to increase scrutiny when a program scores 4.2 out of 5 on attendee satisfaction. But as L&D line items grow, the question shifts from “did people enjoy it” to “what changed.” Satisfaction scores cannot answer that question. Neither can attendance logs or hours-per-employee metrics. The only answer that holds up is behavioral evidence, collected before the program began and again after it ended.
What Kirkpatrick’s Level 3 Actually Requires
The Kirkpatrick four-level framework has been available since 1959. Level 1 is satisfaction. Level 4 is business results. The level that most predicts whether a program produced anything lasting is Level 3: did behavior change on the job? Measuring that level requires structured observation over time, feedback from the manager’s direct supervisor, and assessments conducted before the program and again after it ends, not only at program close. Most organizations never reach Level 3. The data collection is harder, the timeline is longer, and the results are more accountable. A program that scores well on satisfaction but produces no behavior change is far easier to defend when no one has measured behavior in the first place.
A Market-Wide Accountability Gap
No provider visible in current research on measuring leadership training success offers a scored behavioral index measured at three separate time points with a behavior-change guarantee attached. That structure, scoring specific manager behaviors before a program, at its conclusion, and ninety days after, is what separates measurement from the appearance of measurement. Without a baseline, there is no way to attribute change to the program. Without a ninety-day post score, there is no way to know whether any change held. The gap is not about sophistication. It is about accountability, and the market has been avoiding it systematically.
The downstream cost of this avoidance is real. Companies with strong L&D practices achieve 218% higher income per employee and 57% higher retention than peers with weaker programs. Capturing that differential depends entirely on knowing which behaviors changed and which did not. An organization that cannot answer that question is not managing its leadership development investment. It is hoping.
What Measurable Behavior Change Actually Looks Like
The Manager Effectiveness Index (MEI) is built around fifteen specific manager behaviors, organized into five dimensions: clear expectations, coaching conversations, delegation, accountability, and feedback and trust. Each dimension is scored by the manager’s own direct supervisor, not by the manager completing a self-assessment. That distinction matters. Self-report data introduces the exact blind spots that make most training evaluations unreliable. A manager who just completed a program on delegation will almost certainly rate their own delegation skills higher than their boss will. Boss-side scoring removes that distortion and produces a baseline that reflects observable behavior, not good intentions.
Three Measurement Points, One Honest Signal
The MEI is administered three times: before the program begins, at program end, and ninety days after the program concludes. The first score sets the baseline. The second captures immediate learning. The third is the only one that actually answers the question of whether behavior changed on the job. Training transfer, which is the movement of learned behavior from the training environment into real work situations, cannot be confirmed at program end. It can only be confirmed when a manager’s boss scores the same behaviors ninety days later and the numbers have moved. That ninety-day post-measurement is where transfer either shows up or it does not, and most programs never measure it at all.
What a Score Shift Looks Like in Practice
Consider a manager whose boss rates delegation as a consistent gap at baseline, scoring it low across the dimension. The program runs. Eight to twelve weeks later, the end-of-program score is slightly higher. Ninety days after that, with reinforcement coaching completed, the same boss scores delegation higher again, meaningfully and consistently across the dimension’s component behaviors. That is a documented behavior change. It is tied to a specific training investment. It is observable, scored, and recorded at three points in time by the person best positioned to observe it. That is what behavior change evidence looks like when a measurement structure is designed to capture it.
The Manager Performance Cohort and What the Guarantee Actually Means
The Manager Performance Cohort covers one capability, runs eight to twelve weeks virtually, and accommodates up to seven managers per cohort. All three MEI measurement points are built into the program structure, along with ninety days of post-training coaching after the formal program ends. Enterprise cohorts carry the Ninety-Day Behavior Change Guarantee: if MEI scores do not improve and program conditions were met, Tandem runs another coaching cycle at no charge.
The guarantee is a structural response to the satisfaction-score problem described earlier in this post. It is not a confidence statement. It changes who is accountable for what. When a provider’s compensation is tied to a satisfaction survey, the incentive is to deliver content that participants enjoy. When a provider guarantees a score improvement or runs the coaching again for free, the incentive is actual behavior change. The conditions that must be met, including the manager’s participation in coaching and the boss completing all three ratings, are not loopholes. They are the same conditions required for any behavior change intervention to work. The guarantee holds the provider accountable within a fair structure, not an unreachable one.
The Team-Level Case for Measuring Manager Behavior
The downstream effect of manager behavior change extends beyond the individual manager. Research consistently links ongoing manager coaching to team performance: 68% of employees report that their own performance improves when their manager receives ongoing coaching. That figure reframes how the ROI of a program like this should be calculated. A single manager leading a team of seven or eight people means that measured behavior change in one person carries a documented performance effect across the whole team. The investment calculation is not manager-level. It is team-level, and the MEI makes that case with scored data rather than anecdote.
For HR, L&D, and Executive Buyers at Mid-Market and Enterprise Companies
This section is for HR, L&D, and executive buyers at mid-market and enterprise companies. If that is not you, the Academy and association sections below address your situation directly.
The Engagement Gap Has a Price Tag
Global manager engagement sits at 22%, down from 31% in 2022. In best-practice organizations, 79% of managers are engaged. That is not a benchmark to file away. It is a 57-point gap between what most organizations accept and what is demonstrably achievable. Disengaged managers do not coach. They do not develop their people, hold clear expectations, or build the accountability structures that retain top performers. The cost moves through productivity, culture, and attrition simultaneously. Only 38% of employees feel their current leaders are adequately prepared to handle future business challenges. Organizations with mature leadership development programs are 3.5 times more likely to outperform peers in financial results. The gap between where most organizations are and where best-practice organizations operate is not a soft people problem. It carries a measurable financial penalty.
Retention Risk Is Concentrated at the Top of Your Talent Pool
78% of workers cite L&D investment as a deciding factor in joining or staying with an employer. That figure alone justifies the budget conversation. The sharper number is this: 72% of high-potential employees would leave their current organization for one that offers better development opportunities. High-potential employees are the precise group most expensive to replace and most capable of walking out the door with a competing offer in hand. Leadership training functions as a retention lever only when it produces visible, measurable development. Completion rates and satisfaction surveys do not signal to a high performer that the organization is investing in their growth. Behavioral change does. 68% of employees report improved performance when their manager receives ongoing coaching, which means the return on manager development flows downstream to individual contributor performance as well.
CFO Scrutiny Is Arriving at the Same Time as Budget Growth
70% of HR executives plan to increase L&D budgets by 10% or more. U.S. companies spent $102.8 billion on training in 2024 and 2025. Workplace learning data for 2026 confirms that budgets grew 10 to 15% in that period, driven partly by AI skill demands. Budget growth and accountability growth are moving together. Only 30% of organizations are currently effective at using learning data to inform business decisions, according to ATD research. A program measured by satisfaction scores cannot answer the question a CFO will ask when justifying a 10% budget increase: did this make money, save money, or reduce risk? Average ROI for effective leadership programs is estimated at 3 to 5 times the investment. The programs that survive budget scrutiny are the ones that can show what changed, in whom, and when. Programs that cannot answer those questions are exposed in a rising-accountability environment. Training metrics research makes clear that what gets measured gets defended, and what does not get measured gets cut.
What the Manager Performance Cohort Delivers
The Manager Performance Cohort is built for this specific buying context. One capability at a time, up to seven managers per cohort, delivered virtually over eight to twelve weeks. Behavior is scored by each manager’s own boss before the program begins, at program completion, and ninety days later, using the Manager Effectiveness Index across fifteen behaviors and five dimensions. Three measurement points produce a before, during, and after record that answers the CFO question directly. Enterprise cohorts carry the Ninety-Day Behavior Change Guarantee: if scores do not improve and program conditions were met, Tandem runs another coaching cycle at no charge. Pricing is on the cohort page. For organizations whose needs extend beyond a single manager cohort, executive coaching, change management, and culture work are available as bespoke engagements. The leadership coaching segment is projected to exceed $10 billion by 2026, reflecting genuine organizational demand for development that goes deeper than a workshop. The cohort is the entry point. The bespoke work scales from there.
For Managers and Owners at Small and Mid-Size Businesses
This section is for owners and managers at small and mid-size businesses, and for individual managers whose employer will not fund a leadership program. If you are buying for an enterprise cohort, the section above covers your situation.
The Affordability Gap Is Straightforward
A $15,000 enterprise leadership program is not a realistic option for most small business owners or individual managers. The average small company spends $1,091 per learner annually on all training combined, according to Training Magazine’s 2025 Training Industry Report. A single enterprise cohort seat costs more than ten times that figure. The market for leadership development is large and growing, but the money flows almost entirely toward enterprise buyers. That leaves a significant gap for the manager running a team of six who wants real development, not a one-day workshop with a certificate.
What Tandem Academy Delivers
Tandem Academy is built for this gap. The price is $1,000 a year or $99 a month. For that, a member gets nine leadership courses, an AI coach available every week of the year, live group coaching sessions capped at ten seats, and assessments. The catalog value is $17,825. This is not a stripped-down version of enterprise training. It is the same material, structured to work at a self-serve price.
The AI coach addresses the core problem that undermines most training investments. Content without follow-up fades. Baldwin and Ford (1988) documented training transfer at 5 to 10 percent without reinforcement. A manager working through a course on delegation needs support the following week, when a real situation arises, not a PDF to re-read. The AI coach provides that reinforcement between sessions. The manager returns to work with ongoing support, not just notes and good intentions.
Why the Ten-Seat Cap Matters
Live group coaching capped at ten seats is a design choice, not a constraint. A session with ten people is a coaching conversation. A session with fifty is a webinar. At ten seats, every participant can surface a real situation, hear specific feedback, and engage with peers working through the same challenges. Psychological safety and conversational depth require small groups. This format reflects what the learning works and what small-business budgets can actually support.
Retention Is the Business Case
72% of high-potential employees would leave their current employer for better development opportunities elsewhere. For a small business with a team of eight or ten, losing one strong performer is not an abstract HR statistic. It costs roughly 33% of that person’s annual salary to replace them, before accounting for lost productivity and institutional knowledge. Investing $1,000 a year in becoming a more effective manager is also an investment in keeping the people who make the business work. Career development is the number one controllable reason employees cite for leaving in exit interviews. The math on that is not complicated.
For Association Executives Responsible for Non-Dues Revenue and Member Value
This section is for association executives responsible for non-dues revenue, member value, and retention. The enterprise and small-business sections above address different buyers.
The Non-Dues Revenue Problem Is Structural
Member dues once covered the majority of association revenue. That era is over. By 2016, dues represented roughly 30 percent of professional association revenue, down from well above 90 percent in the mid-twentieth century. The pressure has not eased since. Free networking, free industry content, and the general commoditization of information have eroded the perceived value of membership, and the gap between what members pay and what they feel they receive is widening. Associations are now under real pressure to build revenue from the broader commercial network, including suppliers, exhibitors, and prospects, not just dues-paying members. Driving non-dues revenue with eLearning has become a documented strategic priority precisely because digital education scales without proportional cost.
What the Tandem Academy Partnership Looks Like
Associations can offer Tandem Academy to their full network at zero cost and zero delivery work. No curriculum to build, no instructors to hire, no platform to maintain. The association promotes the program; Tandem runs it. On every $1,000 annual membership sold, the association keeps 30 percent, $300 per person per year, on the first purchase and on every renewal. Monetizing member training is increasingly how associations convert their network into a recurring revenue line rather than a one-time transaction.
The arithmetic is straightforward. Five hundred network participants at $300 per person generates $150,000 a year in non-dues revenue. Because the revenue recurs on renewals, the number grows without additional sales effort. Unlike event sponsorships, this model does not reset to zero after each conference.
The Member Value Case
The member value argument is grounded in workforce data. 78 percent of workers say L&D investment is a factor in where they choose to work. 81 percent believe development should be continuous rather than event-based. An association that delivers nine leadership courses, an AI coach available every week of the year, and live group coaching at $1,000 a year is offering something most members cannot source independently at that price. The catalogue value of the Academy is $17,825. Members paying $1,000 are getting enterprise-grade training at a price point that is difficult to find elsewhere.
The Network Effect
The program is not limited to active members. Suppliers, exhibitors, and prospects can all access it, which converts the Academy from a member benefit into a full network engagement tool. Prospects receive real value before they join. Suppliers and exhibitors get a stickier relationship than a booth fee creates. Members who rely on the training have a concrete reason to renew. The revenue line, the recruitment tool, and the retention mechanism are the same program.
The Verdict: Training Is Not the Problem. The Model Is.
Leadership training bought as an event, measured by satisfaction scores, and forgotten ninety days later produces 5 to 10 percent transfer to the job (Baldwin and Ford, 1988; Joyce and Showers). That figure is not a failure of specific programs. It is the predictable output of a broken model, and it is the industry norm.
The fix is structural. Reinforce the learning after the program ends. Measure specific behaviors before training begins, at completion, and ninety days later. Hold the result accountable with something more binding than a feedback form. Those three steps separate programs that change behavior from programs that generate receipts.
The path forward depends on who you are. Enterprise and mid-market HR, L&D, and executive buyers should look at the Manager Performance Cohort. Owners and managers at small and mid-size businesses, and individual managers whose employer will not fund a program, can access the same quality of training through Tandem Academy at $1,000 a year or $99 a month. Association executives looking to generate non-dues revenue while delivering genuine member value should review the association program, which returns 30 percent, or $300 per person per year, on every membership sold.
Global manager engagement sits at 22 percent. In best-practice organizations, that number is 79 percent. The organizations closing that gap are not spending more. They are measuring differently.

