How to Design an AI Roleplay Scenario That Actually Changes Behavior

The software part of building a roleplay scenario takes a few minutes. The part that decides whether it works is a set of design choices most teams make quickly and regret later: how specific the learner's goal is, what the avatar is allowed to withhold, and whether the rubric measures the thing you actually certify on.

Here is the whole anatomy, with a worked example built live on stage.

  September 22, 2026

Key takeaways

Key takeaways
  • Six inputs define a scenario: type, situation, learner goal, avatar, personality traits, and background context. A rubric is selected last.
  • The learner goal is the highest-leverage field. Vague goals produce vague scores, and vague scores produce nothing.
  • Background context is what separates a persona from an answer key. Real colleagues withhold. Your avatar should too.
  • Difficulty should come from realistic resistance, not from an arbitrarily hostile persona. Hostile is easy to dismiss. Reasonable and immovable is hard.
  • Learners see an information panel before starting, covering their role, the avatar's role, and the skills being measured.

The six inputs that define a scenario

At ATD Demo Day, Rapport's Chen Zhang built a complete conflict resolution scenario in front of a live audience, then ran it immediately. The whole build was a handful of fields.

Here is what each one does and what it was set to.

Input
What it controls
Worked example
1. Scenario type
The category, which sets the default rubric and report structure
Growth and development
2. Situation
What is happening, the setting, and why this conversation exists now
A direct report is in conflict with a peer, and a report delivery slipped as a result
3. Learner goal
What success means, and therefore what the rubric measures
As the manager, resolve the conflict
4. Avatar
Who the learner faces, including appearance, accent, and cadence
Anya, the direct report. US English accent, natural cadence
5. Personality traits
Tone, manner, patience, and how the avatar speaks to the learner. Up to three
Defensive, conflict-averse, high-standards
6. Background context
Internal thoughts and feelings the avatar holds but does not volunteer
"Fears she'll solely be blamed if the project fails and feels unsupported by leadership. Believes her high standards are non-negotiable for project success and sees the conflict as a result of the peer's performance, not a shared issue."
Figure 1. Fields one through four are logistics. Fields five and six are where the scenario becomes worth running.
Rapport AI roleplay scenario builder showing the Anya avatar with accent, cadence, personality traits, and background context fields
Inputs four through six in the builder: the avatar, three personality traits, and the background context she will not volunteer.

Once these are set, the scenario is named and saved to a library for reuse and assignment. Nothing here requires a vendor ticket, which is the operational point: a training team that spots a gap on Tuesday can have a scenario live on Wednesday.

An instructional designer planning an AI roleplay scenario with sticky notes at her desk
The build takes minutes. The thinking behind the goal, the persona, and the rubric is where the time should go.

Getting the learner goal right

This is the field that repays the most thought, because everything downstream inherits from it. The rubric scores against what success means, and the learner goal is where success is defined.

The test: could two people read your learner goal and disagree about whether a given conversation achieved it? If yes, the goal is too loose, and the resulting report will be too loose to act on.

Too vague
Specific enough to score
"Have a productive conversation about performance."
"Name the missed deadline, explain its impact on the client review, and secure a commitment to flag future slippage in advance."
"Handle the customer's complaint."
"De-escalate the customer, establish what actually failed, and agree a resolution and a date without promising a refund."
"Pitch the product effectively."
"Uncover the buyer's current workflow, connect two product capabilities to problems they named, and book a technical evaluation."
"Manage up successfully."
"Reset the delivery date with your director, present the trade-off between scope and timeline, and leave with one option agreed."
Figure 2. The right column takes about ninety extra seconds to write and changes every report the scenario will ever generate.

The specific versions also tell you something useful before anyone runs the scenario: if you cannot write the goal, your organization has not decided what good looks like for that conversation, and no amount of practice will fix that.

Background context, the hidden half of a persona

Personality traits control how the avatar sounds. Background context controls what it knows and will not say, and that is the field that decides whether the learner has to work.

Look closely at the demo's background context, quoted in full in Figure 1. It is three sentences and it does an enormous amount of work.

Anya fears being blamed alone. She feels unsupported by leadership.

She treats her own high standards as non-negotiable and reads the conflict as her peer's performance problem rather than a shared one.

She does not open with any of that. It surfaces only when the manager asks about the dynamic, and it reframes the entire conversation when it does.

Note also what it produces behaviorally: a person who fears sole blame will deflect toward a peer, which is exactly what makes this scenario hard to score well on. That is the shape of a real workplace conversation, where the useful information is available but not offered.

Three sentences is about right. Long enough to give the persona a stable internal logic, short enough that it does not contradict itself.

Things worth putting in it:

  • A private diagnosis. What this person actually thinks is wrong, which usually differs from the stated issue.
  • A history the learner may not know. Previous attempts, previous promises, a prior conversation that went badly.
  • A constraint they are protecting. Something they are unwilling to give up and will negotiate hard around.
  • A condition for cooperation. What would have to be true for them to agree, which is what turns a concession into something the learner has to earn. Anya agrees to the meeting only once it is framed as forward-looking.
  • An emotional register that shifts. How they respond if the learner gets it right, and how they respond if the learner pushes too hard.

The avatar responds to the learner's tone of voice as well as their words, in both its replies and its facial expressions, so a persona with a shifting register will visibly react to a rushed or combative delivery. For more on writing personas specifically, see how to write a strong prompt for your AI avatar.

See a well-built persona from the other side

Run the peer conflict scenario from the demo. Anya has a private opinion she will not lead with, and finding it is the exercise.

Try the scenario
A professional listening with a guarded expression during a one-on-one workplace conversation
Real colleagues hold things back. Background context is what gives your avatar something worth uncovering.

Choosing the rubric, and why it should be yours

The grading framework is selected last, but it is doing the heaviest work, because it converts a conversation into something coachable.

Two routes are available. Built-in report templates are based on established communication frameworks, including SBI, and are a reasonable starting point for general management conversations.

The second route matters more for mature programs: organizations with proprietary or specific performance rubrics can upload their own through an enterprise plan and score every learner against that standard.

The reason to use your own rubric is continuity. If your existing training materials reference a named model, and your performance reviews reference that model, and then your practice tool scores against a generic communication template, learners get three different definitions of good.

Aligning the rubric to what you already certify on removes that.

A well-built rubric makes failure specific. In the demo, the report did not say the conversation was weak.

It said two named criteria were unaddressed: describing observable behavior, and explaining impact. That distinction is the subject of why managers score lower than they think, and it is only possible because the rubric named the criteria in advance.

Learners see those criteria before they begin. An information panel covers their role, the avatar's role, and the skills they will be measured against, so the report reads as coaching against a known standard rather than a surprise verdict.

Five ways scenarios go soft

Most scenarios that fail do not fail dramatically. They just turn out to be easy, and easy practice produces confidence without capability.

The mistake
The fix
The cooperative avatar
Give the persona a condition for agreement and let it hold that line until the condition is met.
The hostile avatar
Hostility is easy to dismiss as unrealistic. Reasonable and immovable is harder and truer. Dial resistance, not aggression.
The vague goal
Rewrite until two reviewers would agree on whether a given run achieved it.
The generic rubric
Use the framework your training materials and reviews already reference, uploaded if it is proprietary.
The one-and-done assignment
Assign the scenario as a set of attempts. One run produces a score. Several produce a habit.
Figure 3. The second row catches people out. Teams often equate difficulty with anger, when the hardest conversations involve someone perfectly polite who will not move.

Once a scenario is live, the cohort view tells you whether it is calibrated. If everyone scores well immediately, the scenario is too soft.

If nobody improves across attempts, the rubric may be measuring something the training never taught. Reading those patterns is covered in completion rates are not readiness.

Two L&D colleagues reviewing roleplay scenario performance on a wall-mounted screen
Once a scenario is live, cohort scores tell you whether it is calibrated, too soft, or measuring the wrong thing.

Frequently asked questions

How long does it take to build an AI roleplay scenario?

A few minutes, including persona and rubric. Rapport built a complete conflict resolution scenario live inside a 30-minute conference session.

The real time cost is deciding what good looks like, not operating the software.

What is the most important field in a roleplay scenario?

The learner's goal. It defines success, which determines what the rubric measures.

A goal like "have a good conversation" produces scoring nobody can act on. A goal naming the behavior and the outcome produces a report a learner can use on the next run.

Why does an AI roleplay avatar need background context?

Because real people withhold. Background context holds the internal thoughts and motives the avatar will not volunteer, such as a private belief about who is really at fault.

Without it the avatar answers cooperatively and the learner never has to uncover anything.

Can we use our own scoring rubric for AI roleplay?

Yes. Built-in templates follow established frameworks such as SBI, and organizations with proprietary rubrics can upload their own through an enterprise plan so learners are scored against the standard already used elsewhere in the business.

What makes an AI roleplay scenario too easy?

A cooperative persona, a vague learner goal, and no hidden information. Difficulty should come from realistic resistance rather than from arbitrary hostility, which learners correctly read as fake.

Do learners see the scenario details before they start?

Yes. An information panel covers the learner's role, the avatar's role, and the skills being measured, so the report lands as coaching against a known standard.

Where to go next

The full session this example came from is in our ATD Demo Day recap. For the format fundamentals, see AI role play training and communication role play scenarios that change behavior.

Sales teams designing their own should read AI roleplay software for sales training and diagnose the room before you pitch. Before rollout, run the checklist in AI roleplay and employee data.

Ready-made starting points live in the scenario library, including managing up, delegation, and leading through change.

Build one this week.

Start from a template or write your own persona, set the rubric your team already uses, and have it in front of learners before the next cycle.

Browse the scenario library See pricing