How to Design an AI Roleplay Scenario That Actually Changes Behavior
The software part of building a roleplay scenario takes a few minutes. The part that decides whether it works is a set of design choices most teams make quickly and regret later: how specific the learner's goal is, what the avatar is allowed to withhold, and whether the rubric measures the thing you actually certify on.
Here is the whole anatomy, with a worked example built live on stage.
September 22, 2026
Key takeaways
- Six inputs define a scenario: type, situation, learner goal, avatar, personality traits, and background context. A rubric is selected last.
- The learner goal is the highest-leverage field. Vague goals produce vague scores, and vague scores produce nothing.
- Background context is what separates a persona from an answer key. Real colleagues withhold. Your avatar should too.
- Difficulty should come from realistic resistance, not from an arbitrarily hostile persona. Hostile is easy to dismiss. Reasonable and immovable is hard.
- Learners see an information panel before starting, covering their role, the avatar's role, and the skills being measured.
The six inputs that define a scenario
At ATD Demo Day, Rapport's Chen Zhang built a complete conflict resolution scenario in front of a live audience, then ran it immediately. The whole build was a handful of fields.
Here is what each one does and what it was set to.

Once these are set, the scenario is named and saved to a library for reuse and assignment. Nothing here requires a vendor ticket, which is the operational point: a training team that spots a gap on Tuesday can have a scenario live on Wednesday.

Getting the learner goal right
This is the field that repays the most thought, because everything downstream inherits from it. The rubric scores against what success means, and the learner goal is where success is defined.
The test: could two people read your learner goal and disagree about whether a given conversation achieved it? If yes, the goal is too loose, and the resulting report will be too loose to act on.
The specific versions also tell you something useful before anyone runs the scenario: if you cannot write the goal, your organization has not decided what good looks like for that conversation, and no amount of practice will fix that.
Background context, the hidden half of a persona
Personality traits control how the avatar sounds. Background context controls what it knows and will not say, and that is the field that decides whether the learner has to work.
Look closely at the demo's background context, quoted in full in Figure 1. It is three sentences and it does an enormous amount of work.
Anya fears being blamed alone. She feels unsupported by leadership.
She treats her own high standards as non-negotiable and reads the conflict as her peer's performance problem rather than a shared one.
She does not open with any of that. It surfaces only when the manager asks about the dynamic, and it reframes the entire conversation when it does.
Note also what it produces behaviorally: a person who fears sole blame will deflect toward a peer, which is exactly what makes this scenario hard to score well on. That is the shape of a real workplace conversation, where the useful information is available but not offered.
Three sentences is about right. Long enough to give the persona a stable internal logic, short enough that it does not contradict itself.
Things worth putting in it:
- A private diagnosis. What this person actually thinks is wrong, which usually differs from the stated issue.
- A history the learner may not know. Previous attempts, previous promises, a prior conversation that went badly.
- A constraint they are protecting. Something they are unwilling to give up and will negotiate hard around.
- A condition for cooperation. What would have to be true for them to agree, which is what turns a concession into something the learner has to earn. Anya agrees to the meeting only once it is framed as forward-looking.
- An emotional register that shifts. How they respond if the learner gets it right, and how they respond if the learner pushes too hard.
The avatar responds to the learner's tone of voice as well as their words, in both its replies and its facial expressions, so a persona with a shifting register will visibly react to a rushed or combative delivery. For more on writing personas specifically, see how to write a strong prompt for your AI avatar.
See a well-built persona from the other side
Run the peer conflict scenario from the demo. Anya has a private opinion she will not lead with, and finding it is the exercise.
Try the scenario →
Choosing the rubric, and why it should be yours
The grading framework is selected last, but it is doing the heaviest work, because it converts a conversation into something coachable.
Two routes are available. Built-in report templates are based on established communication frameworks, including SBI, and are a reasonable starting point for general management conversations.
The second route matters more for mature programs: organizations with proprietary or specific performance rubrics can upload their own through an enterprise plan and score every learner against that standard.
The reason to use your own rubric is continuity. If your existing training materials reference a named model, and your performance reviews reference that model, and then your practice tool scores against a generic communication template, learners get three different definitions of good.
Aligning the rubric to what you already certify on removes that.
A well-built rubric makes failure specific. In the demo, the report did not say the conversation was weak.
It said two named criteria were unaddressed: describing observable behavior, and explaining impact. That distinction is the subject of why managers score lower than they think, and it is only possible because the rubric named the criteria in advance.
Learners see those criteria before they begin. An information panel covers their role, the avatar's role, and the skills they will be measured against, so the report reads as coaching against a known standard rather than a surprise verdict.
Five ways scenarios go soft
Most scenarios that fail do not fail dramatically. They just turn out to be easy, and easy practice produces confidence without capability.
Once a scenario is live, the cohort view tells you whether it is calibrated. If everyone scores well immediately, the scenario is too soft.
If nobody improves across attempts, the rubric may be measuring something the training never taught. Reading those patterns is covered in completion rates are not readiness.

Frequently asked questions
How long does it take to build an AI roleplay scenario?
A few minutes, including persona and rubric. Rapport built a complete conflict resolution scenario live inside a 30-minute conference session.
The real time cost is deciding what good looks like, not operating the software.
What is the most important field in a roleplay scenario?
The learner's goal. It defines success, which determines what the rubric measures.
A goal like "have a good conversation" produces scoring nobody can act on. A goal naming the behavior and the outcome produces a report a learner can use on the next run.
Why does an AI roleplay avatar need background context?
Because real people withhold. Background context holds the internal thoughts and motives the avatar will not volunteer, such as a private belief about who is really at fault.
Without it the avatar answers cooperatively and the learner never has to uncover anything.
Can we use our own scoring rubric for AI roleplay?
Yes. Built-in templates follow established frameworks such as SBI, and organizations with proprietary rubrics can upload their own through an enterprise plan so learners are scored against the standard already used elsewhere in the business.
What makes an AI roleplay scenario too easy?
A cooperative persona, a vague learner goal, and no hidden information. Difficulty should come from realistic resistance rather than from arbitrary hostility, which learners correctly read as fake.
Do learners see the scenario details before they start?
Yes. An information panel covers the learner's role, the avatar's role, and the skills being measured, so the report lands as coaching against a known standard.
Where to go next
The full session this example came from is in our ATD Demo Day recap. For the format fundamentals, see AI role play training and communication role play scenarios that change behavior.
Sales teams designing their own should read AI roleplay software for sales training and diagnose the room before you pitch. Before rollout, run the checklist in AI roleplay and employee data.
Ready-made starting points live in the scenario library, including managing up, delegation, and leading through change.
Build one this week.
Start from a template or write your own persona, set the rubric your team already uses, and have it in front of learners before the next cycle.
Browse the scenario library → See pricing