Skip to content

Buyer's Guide

Build vs Buy an AI Coach

The real question is not build versus buy. It is which of five things you are actually building, from a simple answer bot to a coaching platform tuned to your own competency framework and culture. A chatbot is a perfectly good answer if that is what you need. This guide shows the five levels, what each costs to build, what buying gets you, and the one part you cannot build on any timeline.

“Build versus buy” is the wrong question. The right one is which of five things you are building. There is a ladder from a simple answer bot to a coaching platform tuned to your own competency framework and culture, and most products in the market sit somewhere along it. A chatbot is a perfectly good thing to build if a chatbot is what you need. The mistake is shopping for one level while expecting the outcomes of another, or building a Level 5 when a Level 1 would have done.

Two honest through-lines run down the whole page. First, building a coach is a labor and calendar story, not a token story: LLM inference is a rounding error next to the team and the upkeep. Second, and more important, the software is buildable but the measurement is not, at least not on a project timeline. A validated assessment, a calibrated 360, and benchmarked outcomes are earned over years of data or bought. That single fact is what build versus buy really turns on.

A note on transparency: We built Risely, which sits at the top of this ladder, so we would rather you bought ours. We have tried to earn the read by being specific about real build costs, honest about where building is the right call, and clear about the one thing no build can shortcut. If a chatbot is genuinely all you need, this page will tell you to build it.

The AI coaching maturity ladder

Five levels. Cost, effort, and complexity climb as you go up, and so does what the product can actually do. Named tools sit at the ends; the middle is a category, because that is where most point products live.

Basic
cost · effort · complexity →
Level 1

Answer bot

Answers management questions on demand. Reactive, no memory of the person, no measurement.

Buy: ChatGPT, Claude, Copilot
Level 2

Guided assistant

A persona with a scripted flow. Asks questions, remembers the chat, roleplay. No measurement.

Buy: custom GPTs, roleplay tools
Level 3

Structured coach

A skill model, a baseline, progress tracking, nudges. Measures change, on a directional instrument.

Buy: mid-market AI coaches
Level 4

Measured program

Adds 360 feedback, native Slack and Teams, privacy, and org dashboards. A generic framework.

Buy: enterprise AI coaching platforms
Level 5

Integrated platform

Runs on your own competency framework, content, and culture. Benchmarked outcomes, earned over time.

Buy: Risely

The line that build versus buy really turns on

Every level mixes two very different kinds of work. One is buildable on a timeline. The other is not, no matter the budget. Keeping them separate is the whole trick to reading this decision honestly.

Software capability, buildable

The chat, the persona, memory, nudges, native Slack and Teams, dashboards, content ingestion. This is real engineering, and it is a bounded project. You can reach a first version you are happy with and run it. The cost is a team and time, and the numbers below are defensible.

Nature: money and calendar. Buildable.

Coaching science, earned

A methodology grounded in organizational psychology, a knowledge graph the agent reasons over, a validated assessment, a calibrated 360, and benchmarked outcomes. These are not features you ship; they are the product of years of research and data across large populations. A build reaches a directional baseline quickly, but the validated, benchmarked core is earned over time or bought from someone who already did the work.

Nature: time and data. Not buildable on a timeline.

Read the ladder with this in mind. Levels 1 and 2 are almost entirely the buildable kind, which is why building there is cheap and sensible. From Level 3 up, the words that matter most, “measured,” “calibrated,” “benchmarked,” live on the right-hand side. That is the part buying exists to provide.

Each level: what you buy, what you build

Build figures are US-loaded team cost to a first version, for the software only. Offshore-blended teams run about 55 to 65 percent of these, not the 20 to 40 percent sticker rates imply, once management and rework are priced in. Inference and infrastructure are shown separately because they behave differently.

Level 1

Answer bot

If you buy

A general assistant answers management questions today, for roughly nothing. This is the honest recommendation if on-demand answers are all you want.

Examples: ChatGPT, Claude, Microsoft Copilot · $0 to $30/user/mo

If you build
$5k–15k
1 engineer, 1–2 weeks
$0.20–3
API /user/mo · $10–30k/yr upkeep

A prompt over an API. Genuinely a weekend. If this is your intent, build it and move on.

Level 2

Guided assistant

If you buy

Custom GPTs and point roleplay tools give you a coaching persona and scripted practice without building. Fine for structured practice, not for measured development.

Examples: custom GPTs, AI roleplay apps

If you build
$50k–140k
1–2 eng, 2–4 months
$2–8
API /user/mo · $25–70k/yr upkeep

Prompt engineering, a thin app, and conversation memory. The moment it holds personal conversations, you also inherit an ongoing safety obligation.

Level 3

Structured coach

If you buy

Mid-market AI coaching platforms give you a skill framework, assessments, and progress tracking out of the box, with measurement that has been used across many organizations.

Examples: mid-market AI coaches

If you build
$250k–600k
2–3 eng + fractional I/O, 6–10 mo
$3–10
API /user/mo · $90–220k/yr upkeep

What you get is a directional, unvalidated baseline. It measures something, but it is not a psychometrically validated instrument. That takes its own program, below.

Level 4

Measured program

If you buy

Enterprise AI coaching platforms add 360 feedback, native Slack and Teams, privacy architecture, and org dashboards, tested across many deployments and compliance reviews.

Examples: enterprise coaching platforms

If you build
$700k–1.6M
3–5 eng + I/O + security, 12–20 mo
$4–12
API /user/mo · $180–400k/yr upkeep

Native Teams app development, 360 synthesis, and privacy work are heavy. A calibrated 360 is a genuine psychometric project, not a config step. This is where “2 to 3 engineers” stops being realistic.

Level 5

Integrated platform

If you buy

A Level 5 platform brings the part a build cannot: a coaching agent backed by organizational psychology and a real methodology, running on its own knowledge graph of skills and behaviors, with measurement already validated and benchmarked across a large population. Your own competency framework, learning content, and culture layer on top of that earned core, rather than being built from zero.

Example: Risely, whose Merlin agent runs an 83-skill model calibrated by 360 feedback, with outcomes measured across 5,000+ users and 40+ organizations. From $59/user/month, or $700 to $1,000/user/year at enterprise scale.

If you build
$1.2M–2.6M
4–6 eng + ML/data + security, 15–24 mo
$8–20
API /user/mo · $250–550k/yr upkeep

That is the software, and it is buildable. The earned core is not: the methodology, the knowledge graph, and the validated, benchmarked measurement all rest on research and data you have not collected yet. You start that flywheel at zero on launch day. That, not the code, is what a build cannot buy back.

Add infrastructure and tooling of about $30k to $120k a year (vector database, hosting, observability, evals), separate from the API figure. Assumptions: US-loaded senior engineer roughly $220–300k/year; I/O psychologist roughly $130–180k/year, usually fractional; inference priced with prompt caching and cheap-model routing for nudges, which most competent teams use. Figures are planning estimates, not quotes.

Building a coach is a labor story, not a token story

The number people fixate on when they say “building is cheap now” is inference, and it is the one number that barely matters. At a mid-ladder level, a heavily engaged user costs a dollar or two a month in tokens. One engineer costs a quarter of a million dollars a year. The cost of a build is people and calendar; the API line is a rounding error.

One year, 500 internal users, Level 4
Engineering team (build + maintain)~$900k
Infrastructure and tooling~$75k
LLM inference (500 users × $8 × 12)~$48k

Illustrative. The point is the ratio: inference is about 5 percent of what a real build costs. Anyone selling “just wrap an API” is pricing the smallest line.

20–30%
of build cost per year, ongoing, just to keep an LLM product current as models deprecate and prompts drift. Higher maintenance than normal software.
55–65%
of US cost for an offshore-blended team, once management, ramp-up, and rework are priced in. Not the 20 to 40 percent that headline rates suggest.
~5%
MIT’s 2025 review found only about 5 percent of enterprise GenAI initiatives reached production value, and internally built tools fared worse than bought ones.

What a build cannot buy back

You can ship the software. Beneath it sits the part that took the credible platforms years, and that no build reaches on a timeline. This is the honest reason buying exists, and the strongest single argument on this page.

A methodology backed by org psychology

The logic of a coaching session, grounded in organizational psychology: what to ask, when to challenge, when to hold someone to a commitment. It is what makes the agent coach rather than answer, and it is designed and encoded by people who understand coaching, not improvised in a prompt.

Its own knowledge graph

A structured model of skills, behaviors, and how they connect, that the agent reasons over to personalize coaching. Your learning content layers onto it. This is the difference between an agent that understands the domain and one that merely retrieves documents.

A validated assessment

Not a checklist that looks reasonable, but an instrument with tested reliability and validity, normed on a representative sample that typically runs into the thousands of respondents. Industry validation studies run into six figures per instrument and 12-plus months. A build reaches a directional baseline; a validated instrument is a separate, time-gated program.

A calibrated 360

One of the hardest instruments in the field. Individual rater effects can account for more variance than the trait being measured, so a defensible 360 is a genuine psychometric research effort, not a feature you configure. Getting it wrong produces confident numbers that mean nothing.

Benchmarked outcomes

You cannot benchmark data you have not collected. Credible outcome evidence comes from years of usage across large populations, which is why the established platforms have outcome studies and a fresh build does not. Public frameworks like O*NET shortcut the taxonomy, not the benchmark. This is the flywheel that starts the day you launch, not before.

Ongoing safety and evaluation

The moment conversations turn personal, you own a continuous obligation: sensitive disclosures, crisis language, and a live regulatory surface around AI and wellbeing. This is not a launch milestone; it is standing operating cost, and it starts at the guided-assistant level, not at the top.

One more honest note: the most credible platforms in this market are not pure AI. They pair AI with human coaches and years of accumulated outcome data. A pure-build AI coach is competing with that, from zero data, on day one.

When building genuinely makes sense

Plenty of cases. Pretending otherwise would insult a technical buyer. Match the case to the level.

You only need Level 1 or 2

You want on-demand answers or scripted practice, not measured development. Building is cheap, fast, and the right call. Do not buy a coaching platform for a job a chatbot does.

Coaching is your core product

If you sell coaching or people development, the platform is your moat, not an internal tool, and the build is your R&D. The measurement flywheel is worth starting because it is your business.

Constraints no vendor can meet

Sovereign or air-gapped data residency, or a domain so unusual that no platform fits. Rare, but real, and a legitimate reason to own the stack even at the top of the ladder.

You can staff the science and the calendar

You have engineering, I/O psychology, and security capability, an HR owner to lead it functionally, and the patience for a multi-quarter build plus the years measurement credibility takes. If you cannot staff all of that, buy.

If you need measured development at Levels 3 to 5 and none of the build cases fit, buying wins on cost, speed, and the one thing that matters most: credible measurement you can defend, from day one.

If you decide to build

We would rather help you build it well than watch you relearn it

If you have a genuine reason to build, especially at the top of the ladder, we are not going to pretend it cannot be done. We do it every day. What we can do is save you the expensive part of the learning curve. We have spent years and 5,000+ users working out the skill model, the assessment design, the coaching methodology, the measurement, and the evaluation infrastructure that keeps an AI coach safe and consistent, and we will share that directly.

In practice, most conversations that begin with “we want to build” end somewhere more useful than pure build or pure buy: buy the platform for the coaching program, and let us help you integrate your own competency framework, content, and systems on top of it. You get the outcome and the measurement credibility without owning the product. Either way, we will give you an honest read on which level you actually need and which path fits, including the cases where we would tell you to build.

Transparency

What a Level 5 platform actually brings

The software at the top of the ladder is buildable. What is not, on any timeline, is the earned core beneath it: a coaching methodology grounded in organizational psychology, a knowledge graph of skills and behaviors the agent reasons over, and measurement validated and benchmarked across a large population. Risely is our worked example of that level. Merlin runs on that core, and your own competency framework and content sit on top of it, rather than being built from nothing.

The numbers below are not a pitch. They are the benchmark a fresh build starts without, and that gap, measured in years of data, is the honest heart of build versus buy.

26%

average skill improvement in 12 weeks

83

workplace skills assessed and coached

5,000+

users coached across 40+ organizations

40

languages supported, voice and chat

See the top of the ladder in action

Before you scope a build, spend 14 days inside a finished one. Try Merlin, Risely's AI coach, free. No credit card, no sales call. It is the fastest way to see what separates a chatbot from a coaching program.

Frequently Asked Questions

Should we build our own AI coach or buy one?

It depends which of five levels you actually need. If you want a simple answer bot for management questions, building is cheap and buying is basically free, and that is a fine choice. As you move up the ladder to a structured coach, a measured program, and finally a platform tuned to your own competency framework and culture, building gets expensive fast and one thing stops being buildable at all: validated, benchmarked measurement, which is earned over years of data or bought. Most organizations that need real, measured development should buy, because that is exactly the part a build cannot shortcut.

How much does it cost to build an AI coaching platform?

By level. A basic answer bot is $5k to $15k and a week or two. A guided practice assistant is $50k to $140k over a few months. A structured coach with a skill model and progress tracking is $250k to $600k over 6 to 10 months. A measured program with 360 feedback, native Slack and Teams, and dashboards is $700k to $1.6M over 12 to 20 months. A platform tuned to your own framework and content is $1.2M to $2.6M over 15 to 24 months. Add $30k to $120k a year for infrastructure and, for anything with an LLM in it, ongoing maintenance of roughly 20 to 30 percent of build cost per year, because models deprecate and prompts need re-testing. Offshore-blended teams run about 55 to 65 percent of US cost, not the 20 to 40 percent that sticker rates suggest.

Isn't building cheaper now that LLMs are commoditized?

The model is cheaper, and at the low end that makes building an answer bot genuinely cheap. But the cost of a real coaching program was never mostly tokens. Inference runs about $2 to $20 per user per month even at the top of the ladder, a rounding error next to a team of engineers and the ongoing maintenance an LLM product demands. Building a coach is a labor and calendar story, not a token story, and commoditized models do nothing for the parts that take time: a validated skill model, a calibrated 360, and benchmarked outcomes.

Can we build a validated assessment as part of the project?

Not on the timelines above. Building software is bounded; validating measurement is not. A defensible, validated instrument is its own program: item development, reliability and validity testing, and a representative norming sample that typically runs into the thousands of respondents. Industry validation studies alone commonly run into the six figures and 12-plus months for a single instrument. So a build reaches a directional, unvalidated baseline quickly, but a genuinely validated assessment is earned over time or licensed from someone who already did the work. Do not confuse the two.

What can you never build on a timeline, no matter the budget?

Benchmarked outcome data. You cannot benchmark results you have not yet collected. The platforms whose measurement is credible got there through years of usage across large populations, which is why their outcome studies exist and a fresh build's do not. You can license public frameworks like O*NET for the taxonomy, but norms, validity, and an outcome benchmark are a data-accumulation flywheel that starts the day you launch, not before. This is the honest core of build versus buy: you can build the software, but the measurement credibility either takes years to earn or comes with the platform you buy.

Is a self-built AI coach safe to use for talent decisions?

Treat a self-built coach as a private development and practice tool, not a system of record for measurement or talent decisions. An in-house, unvalidated instrument used in decisions carries real defensibility and adverse-impact risk. There is also a continuous safety obligation the moment conversations turn personal: sensitive disclosures, crisis language, and a live regulatory surface around AI and wellbeing. That safety and evaluation work is ongoing operating cost from the first guided-conversation level upward, not a one-time build task.

Can Risely help us build our own internal AI coach?

Yes. If you have a genuine reason to build, especially at the top of the ladder, we would rather help you do it well than watch you spend a year rediscovering what took us years and 5,000+ users to work out. We can advise on the skill model, assessment design, coaching methodology, measurement, and the evaluation infrastructure that keeps an AI coach safe and consistent. Most conversations that start with build end with a hybrid: buy the platform for the coaching program, and we help you integrate your own competency framework, content, and systems. Book a call and we will give you an honest read on which level you actually need and which path fits.