Almost 60% of first-time managers say they never received any training when they moved into the role. The number comes from Bill Gentry at the Center for Creative Leadership, and it sets the baseline that any development plan is really competing against. Nothing at all.
Every vendor page answers this the same way. The plan is personalized, they say, and then personalization gets described as a loop: collect data, generate insight, tailor guidance, repeat.
The loop is real and it gives a buyer nothing to check. What a buyer can check is the output: which topic comes up in week 4, why that one instead of another, and what would have had to happen for it to be a different topic.
What follows is one AI coaching plan over twelve weeks, including the week it stopped being the plan it started as.
How do AI coaching platforms personalize leadership development plans?
I can only answer that for the one I built. Risely takes a stated development gap, maps it to specific skills at the person’s role and level, and builds a week-by-week program of topics, coaching session agendas, reflection prompts and take-home exercises around that gap, then adjusts the program each week from what surfaced in the last session. Other platforms in this category describe personalization in their own terms, and I have not audited how any of them generates a plan, so read what follows as one worked example rather than a survey of the field.
The part worth examining is how small the input is allowed to be.
On Risely, it’s a sentence. A manager, an HR business partner or the person themselves writes what needs work in plain language, and a request as thin as “needs help with delegation” is enough to generate a plan. There’s no competency framework to configure first and no taxonomy exercise to fund before anyone gets coached. That’s the difference between this and the version built by hand, where writing a personalized learning plan for one person costs a skilled L&D partner real hours, which is why most managers get a generic curriculum instead.
The request can come from the manager rather than from HR, and that’s the right order. Sydney Finkelstein made the point in his 2019 HBR piece on why one-size-fits-all development fails: exceptional bosses don’t leave it to HR to create career progression programs for their team members. A plain-language starting point is what lets a manager act on that without an L&D project behind them.
Two things then get established before week 1. A baseline assessment of the skill, with 360 feedback from the people who actually work with the person. And a mapping of the stated gap onto Risely’s skill library, resolved against that person’s role and level.
Personalization pulls from the organization too. A company can load its own leadership frameworks, manager playbooks and learning content, and tune the coaching to its culture and values, so the plan speaks in the language the manager already hears in performance reviews.
This is also where AI coaching separates from a general assistant. A chat tool can answer a good question about delegation. It has no baseline to measure against and nothing that decides what next week looks like, which we’ve broken down in full in our comparison with ChatGPT.
Choosing what a plan works on stays a real decision even when the input is one sentence. Gallup’s Beck and Harter reported in 2014 that about one in ten people have the talent to manage, which means the person writing that sentence is usually describing someone learning the job rather than someone with a knack for it. There is more than one thing they could be working on, and a plan that starts on the wrong one still runs for twelve weeks.
Take Emma, a first-time manager who needs help with delegation
Take a first-time manager, call her Emma, promoted from senior engineer three months ago. She’s illustrative, not a customer.
Her manager saw the pattern that shows up in promotions out of a technical role. Emma’s team was shipping, and Emma was doing too much of the shipping herself.
So her manager wrote one line into Merlin: Emma needs help delegating to her team instead of absorbing the work herself.
That sentence was the entire input. It produced a twelve-week leadership development plan with a baseline assessment, weekly topics, session agendas, reflection prompts and a second assessment at the end. What follows is what the plan did with it.
Twelve weeks, and what changed after week 3
Emma is illustrative. The mechanics below are Risely’s, and they’re what decide whether a plan is personalized or just scheduled.
| Weeks | What Merlin does | What Emma does | What shifts |
|---|---|---|---|
| 1 to 3 | Runs the baseline assessment with 360 input, sets the opening topic: delegating the task versus delegating the decision | Takes the assessment, has a first coaching conversation about a real handoff already on her calendar | Nothing yet. The plan is still the plan it was generated as |
| 3 | Picks up the migration-plan incident out of the session and re-plans the topics that follow it | Brings the incident itself, rather than a general worry about trust | The next topic moves from choosing what to delegate to agreeing what “done” means before the handoff |
| 4 to 8 | Sequences topics on that new thread and checks the last session’s commitment against the next real situation | Names the gap, commits to a specific action, reports back on what happened | Awareness turns into a commitment, and the commitment gets checked against something that actually occurred |
| 8 to 12 | Works the sessions toward the stated objective, then re-runs the assessment with 360 feedback | Runs the handoffs she’s been practicing, on her own work | The before-and-after comparison, with her team’s input, is what says whether the change held |
Weeks 1 to 3 look like every other plan, and they should. The assessment establishes where Emma actually is rather than where her manager guessed she was, and the first topic draws the line between handing someone a task and handing them the decision inside it. Her first coaching conversation was about a handoff that was already on her calendar for that week.
Then week 3 happened. Emma had asked a senior engineer to own the migration plan, and the night before the review she rewrote his draft. She brought that to the session.
A fixed curriculum would have moved to week 4’s module anyway. This plan didn’t. The next topic changed from choosing what to delegate to agreeing what “done” looks like before the handoff, because the handoff itself had gone fine. What nobody had agreed on was the standard the migration plan had to meet.
That’s the second mechanic, and the one that’s hardest to see from a vendor page: next week’s agenda changes from what came up in the last session. Not a fixed curriculum with optional modules bolted on. The plan re-sequences around whatever the manager actually brought, which also means a session where nothing specific gets described leaves the sequence roughly where it was.
Emma’s manager never filled in a framework or ranked her against a leadership model. He described a problem in a sentence, the way he would have described it to a colleague. That’s the first mechanic, and it had already done its work before week 1: the input is a plain-language request, not a competency intake.
What the topics actually contain is decided by the third mechanic. Skills are mapped to the person’s role and level, across 83 skills and more than 1,000 O*NET occupations. So “delegation” for a first-time manager three months out of an engineering seat resolves to a different set of behaviors than “delegation” for a director handing work to other managers. Same skill name, different plan. A director working on that same named skill would have started with which decisions to push down a layer and which ones to keep. Emma’s plan never went near that question, because it isn’t where she was standing.
Weeks 4 to 8 ran on the thread week 3 opened. The shape of each session was the same: Emma named the gap, committed to one specific action, and the next session opened by checking that commitment against whatever real situation had come up since. Awareness, then commitment, then accountability. The accountability step is the one that needs a next week to exist, which is why a plan with dates beats a library with everything in it.
By weeks 8 to 12 the sessions were working toward the objective her manager had stated in that first sentence. What changes by the end is smaller than the word transformation implies: Emma was agreeing what “done” meant inside the handoff conversation itself, rather than discovering the gap the night before a deadline. Her second assessment, with her team’s input, is what tells her whether the change held. Her own sense of it isn’t enough, and neither is mine.
A fair question at this point is whether spacing the work over twelve weeks does anything that a good two-day workshop wouldn’t. The closest peer-reviewed comparison is old and imperfect.
Grant (2007), in Industrial and Commercial Training, compared a 13-week coaching-skills course of 23 people against a two-day “Manager as Coach” program of 20. The 13-week group improved on both goal-focused coaching skills and emotional intelligence. The two-day group improved only on coaching skills, and by less.
Read that honestly. Those were people learning to coach, not managers being coached, the samples are small, and the study is from 2007. It’s an analogy for spaced versus intensive learning and nothing stronger than that. The weight of this section rests on the three mechanics above, not on Grant.
How does AI coaching help new managers improve leadership skills?
It gives a new manager a place to think through a real situation from their own week before they walk into it, then holds them to what they said they’d do when the next session opens. The situations are theirs, which is the part a workshop can’t copy.
The conversation runs by text or voice in 40 languages, and Merlin runs natively inside Slack and Microsoft Teams, so the coaching happens where the manager is already working rather than in a portal they have to remember to open. Between sessions, daily nudges arrive by email, Slack or Teams. Those nudges do more work than they look like they should: 73% show high engagement with daily nudges, and the days between conversations are where a commitment either gets tested on a real person or quietly expires.
The model underneath is awareness, then commitment, then accountability. A manager has to see the gap before they’ll act on it, name a specific action rather than an intention, and then be asked about it by something that remembers what they said. Risely users average 26% skill improvement in 12 weeks, measured by the before-and-after assessment with 360 input rather than by how many sessions they completed.
If what you want is a week-by-week checklist for the first three months in the job, we’ve written that separately as a 90-day leadership plan for new managers. If you’re looking at this as a program across a population rather than for one manager, the program-level view of AI and leadership development is the better starting point.
What the plan can’t see
The plan personalizes around what gets named or assessed. A blind spot nobody raised stays invisible to it. A human coach sitting in on one of Emma’s team meetings might catch something that no assessment or plain-language request was ever going to surface.
Merlin doesn’t read meetings, calendars or messages. The plan adapts to what the manager brings to it, so a manager who under-reports, or who skips the practice and says it went fine, gets a plan built on that version instead. None of that is independently verified.
Memory is worth stating precisely, because it’s often oversold. Merlin remembers across sessions and builds on earlier conversations, and what it knows comes only from those sessions.
Where a human coach is the right answer, this doesn’t substitute for one, and we’ve laid out where each fits in our comparison of AI and human coaching.
Then there’s who sees what, which is worth settling before a pilot rather than after one. On an assigned plan, the organization sees session summaries, engagement and before-and-after skill scores, and never the conversation content. Self-driven coaching is fully private. That boundary is what makes the honesty in a session possible, and it also means a manager who wants to hide a struggling quarter from their employer can, because the coaching they’d most benefit from is the coaching nobody will ever see.
What moved Emma wasn’t the personalization engine. It was one mundane handoff becoming a specific committed action, checked against whatever real situation came next. Nothing about the migration plan was dramatic. She rewrote a colleague’s draft late at night, the way plenty of managers do without ever mentioning it, and the only unusual part is that she described it out loud in a week when something was listening and able to act on it.
Personalization earns its keep only because it keeps that loop pointed at whatever is actually happening this week.
This is judgment, not data: a coach changes the next session when someone describes a specific incident rather than a general worry, because the incident is the material. The re-planning step is that judgment written down and run every week.
So when you’re evaluating platforms, don’t ask whether a vendor can generate a plan. All of them can. Ask what happens to week 6 when week 5 goes badly, and make them show you the mechanism.
Try Merlin free with a situation from your own week.
