We Built an AI Roleplay Live at ATD Demo Day. The Manager Scored 45.

Rapport closed out ATD Demo Day with a session that had no slides after the first three minutes. Chen Zhang built a conflict resolution scenario from scratch, ran it live against an avatar who pushed back, and then put her own scorecard on screen in front of a few hundred L&D leaders.

She scored 45. That number is the reason this recap is worth reading.

  September 22, 2026

Key takeaways

Key takeaways
  • A complete roleplay scenario, including persona, hidden motivations, and a grading rubric, was built in a few minutes with no vendor ticket.
  • The conversation ended with both parties agreeing to a meeting. It still scored 45, because agreement is not the same as a well-run feedback conversation.
  • The two criteria that sank the score were describing observable behavior and explaining impact. Neither is a knowledge gap. Both are habits.
  • Individual reports stay with the learner by default. Aggregate scores roll up to a cohort dashboard so trainers can see which competency to coach next.
  • On data: no training of third-party AI models on customer content, audio and full transcripts are not retained long term, and servers are US-based with EU and Asia options for enterprise.

View the webinar

The gap we came to talk about

Chen opened with the problem that makes AI role play training worth building at all: the expensive distance between what a team knows and what it can do under pressure.

Her example was deliberately ordinary. A manager can have every conflict resolution framework memorized, a performance review template ready, and talking points prepared.

Then an upset direct report gets emotional, and all of it goes out the window. Knowledge was never the constraint.

Composure under stakes was.

That framing sets up three problems that conventional training does not solve, and each one shaped a different part of the demo:

  • Practice without stakes does not transfer. When pushback feels real, skills develop fast. When it feels like a script, nothing sticks. This is the same argument behind why most skills development programs don't stick.
  • Readiness is invisible and subjective. Managers and trainers are asked to certify soft skills they have no objective read on, which is the core of what your LMS can't show you about soft skills.
  • Great coaching does not scale. Almost nobody gets a mentor sitting beside them in every hard conversation. Simulation is how that coaching reaches a distributed team.
640,000+
role plays run on the Rapport platform to date, across many languages and dialects
45
score on a conversation that ended in agreement and felt successful to the manager
2
scorecard criteria that accounted for most of the gap, both conversation habits

Building the scenario live, in a few minutes

The first segment was scenario authoring. The point being made was about effort: this is a task a training team does itself, in minutes, not a services engagement.

Six inputs produced a complete, reusable scenario:

Input
What it sets
What was used in the demo
Scenario type
The category, which shapes the default rubric and report
Growth and development
Situation
Plain-language description of what is happening and why this conversation exists
A direct report is in conflict with a peer, and delivery has slipped
Learner goal
What the learner is trying to achieve. Chen called this the most important field
Resolve the conflict as the manager
Avatar
Who the learner faces, including appearance and voice
Anya, the direct report
Personality traits
Tone, manner, and how the avatar speaks to the learner
High standards, frustrated, quick to assign blame elsewhere
Background context
Internal thoughts and feelings the avatar does not volunteer
Private belief that the peer's working style is the real problem
Figure 1. The background context field is the one people underestimate. Real colleagues hold things back, and a persona that does the same is what makes the conversation worth having.
Rapport AI roleplay scenario builder showing the Anya avatar with accent, cadence, personality traits, and background context fields
Setting up the avatar in the scenario builder: name, accent, cadence, up to three personality traits, and the private background context that shapes how she responds.

A grading framework is selected last. Built-in options are based on established communication frameworks, and organizations with their own rubrics can upload them.

If you want the deeper version of this process, we broke it down in how to design an AI roleplay scenario that actually changes behavior, and how to write a strong prompt for your AI avatar covers persona wording specifically.

The conversation with Anya

Chen then ran the scenario live as Anya's manager, on a Zoom stream, without a net. The exchange lasted roughly two minutes and went the way these conversations usually go.

Anya opened by attributing the delay to her peer, Taylor. Chen acknowledged the frustration and asked about the dynamic.

Anya described a personality clash and said her efforts to set expectations just bounce off. Chen proposed a three-way facilitated conversation.

Anya pushed back that it would become a blame session, and only agreed once Chen committed to keeping it forward-looking and focused on workflow. They settled on a time for the three-way meeting.

OK, so it sounded like pretty good. Upon first blush, I'm pretty happy with this, but I have a sense that the report is going to tell me there's some areas I could have improved as well. Chen Zhang, Chief Growth Officer at Rapport, immediately after the live roleplay

This is the part worth sitting with. The avatar did not melt down or stonewall.

It negotiated, held a position, conceded on specific terms, and closed cooperatively. By the usual standard, which is whether the meeting ended well, the manager won.

Run this exact scenario yourself

Peer conflict, an avatar with a private opinion she will not lead with, and a scored report at the end. It takes a few minutes and no setup.

Try the peer conflict scenario

The scorecard: 45

The report came back at 45, which visibly surprised the room and, briefly, the presenter. The top-level view showed partial credit on opening the context and describing the situation.

Two criteria were effectively unaddressed.

Criterion
Result
What the report explained
Open the context
Partial credit
The conversation was framed and the purpose was stated, so the opening did its job.
Describe the situation
Partial credit
The topic was raised, but at the level of "that report delivery" rather than anything specific.
Describe observable behavior
Not addressed
No specific action was named. The manager went straight from "this happened" to "what do we do about Taylor," so the direct report never learned which behavior was the issue.
Explain the impact
Not addressed
The consequences of the missed timeline were never stated. Without impact, the conversation gives the employee no reason to weigh the change as important.
Figure 2. Both failures are structural, not informational. The manager knew the situation perfectly well and simply did not say the two things that make feedback land.
Rapport AI roleplay learner scorecard showing an overall score of 45 percent with no credit for describing observable behavior or explaining the impact
The learner report: 45 overall, strong marks for opening, dialogue, and next steps, and empty bars for observable behavior and impact.

That pattern is not unusual, and it is the subject of its own breakdown in why managers score lower than they think. Describing observable behavior is the gold standard for feedback, and explaining impact is what turns an observation into a reason to change.

Skip both and you get exactly what happened here: a pleasant conversation, a scheduled meeting, and an employee who still does not know what she did.

The learner report closes by telling you what to do on the next run, which matters more than the score. The loop is short enough that a learner can immediately go again and apply the correction, which is the mechanism the science of engaging soft skills training covers in depth.

From one score to a cohort picture

The last segment swapped the learner hat for the leadership one. Every learner's per-criterion scores roll up into a dashboard that answers three questions: how is the team performing overall, how are they doing on one specific scenario, and how is one specific learner progressing.

Chen showed performance by metric across two scenarios, then drilled into an individual. The useful move is comparing scenarios, because a competency that holds up in one context and collapses in another tells you something a single average never will.

We go further into reading these views in completion rates are not readiness and how to run a skill gap analysis.

In most deployments the roleplay itself is embedded inside the customer's LMS, so learners reach it through the system they already use, while managers log into Rapport for the analysis layer. Rapport is SCORM compliant, which is what makes that arrangement straightforward.

Rapport Cohort Analytics dashboard showing average overall score, key areas for improvement, highest-scoring competencies, and performance by metric
The cohort view turns every learner report into team-level signals: the weakest competencies, the strongest ones, and performance by metric across scenarios.

What the room actually asked

Eliza Quigley ran the chat throughout, and the questions clustered in three places. They are a good proxy for what any L&D team will ask internally.

Data and security

This was the largest cluster by far, and it gets its own full treatment in AI roleplay and employee data. The short version given on the call:

  • A secure knowledge base holds your content, rubrics, and scenarios, and is used only to drive your scenarios and reports.
  • No data sharing with AI models by default. Customer conversations are not used to train AI models.
  • Audio and explicit transcripts are not stored long term. They are processed to produce the report, and what persists is login information, user ID, and high-level scoring outputs.
  • Servers are US-based, with EU or Asia instances available for enterprise customers. FedRAMP authorization is under consideration, so teams with that requirement should confirm current status with Rapport.

How the avatars behave

Avatars range from photorealistic to stylized characters, and the visual can be turned off for audio-only practice. The avatar responds to the learner's tone of voice in both its replies and its facial expressions.

Learner body language is not an input today. Before starting, learners see an information panel covering their role, the avatar's role, and the skills they will be measured against, so nobody walks in blind.

Fitting it into an existing program

Scenarios can be embedded in an LMS, shared as a private link, or accessed through the Rapport learner portal with single sign-on. Results funnel back to the same cohort view regardless.

Trainers author scenarios and assign them; learners run them. Teams can start from templates, including managing up and de-escalation, or have Rapport build an initial knowledge base from existing training materials and then take over authoring.

See learning and development solutions and integrations for the mechanics.

Frequently asked questions

How long does it take to build an AI roleplay scenario?

A few minutes. The trainer picks a scenario type, describes the situation, defines the learner and avatar, sets the learner's goal, adds personality traits and background context, and selects a grading framework.

The scenario then saves to a reusable library.

Who creates AI roleplay scenarios, trainers or learners?

The training team creates and assigns scenarios. Learners run what they are assigned rather than authoring their own.

Rapport can build the first set from your existing materials, after which your team takes over.

Does an AI roleplay avatar respond to body language?

Not today. The avatar does respond to the learner's tone of voice, in both its replies and its facial expressions.

Avatars range from photorealistic to stylized, and the visual can be switched off for audio-only practice.

Can AI roleplay use our own scoring rubric?

Yes. Built-in report templates are based on established communication frameworks, including SBI.

Organizations with proprietary rubrics can upload them through an enterprise plan and score every learner against their own standard.

Does AI roleplay need to be embedded in an LMS?

No. Rapport is SCORM compatible and embeds cleanly, which helps adoption since learners are already logged in.

Teams without an LMS can share a private scenario link or use the learner portal with single sign-on.

Results roll up to the cohort view either way.

Do managers see every learner's roleplay report?

Not by default. Individual reports stay with the learner, because psychological safety is part of the design.

Aggregate scores roll up to the dashboard, and learners can export their own report to bring to a coaching session. Enterprise customers with a specific business need can arrange broader access.

Where to go next

If you are new to the format, start with the AI role play training pillar. If you are comparing vendors, the best AI roleplay platforms covers the field and AI roleplay vs manager roleplay covers when simulation is the wrong tool.

Managers focused on the conversations themselves should read communication role play scenarios that change behavior. Teams building a business case will want the business cost of skipping soft skills.

Find out what your managers would score.

Run the same conflict scenario Chen ran on stage, get the report, and see which two criteria your own conversation leaves on the table.

Try the scenario free Book a demo