Why Managers Score Lower Than They Think: The Two Things Missing From Most Feedback Conversations

At ATD Demo Day, a seasoned executive ran a live conflict resolution roleplay in front of a few hundred L&D leaders. The conversation ended in agreement.

A meeting got scheduled. Everyone was polite.

The scorecard came back at 45%, and the report named the reason in three words: "Impact not explained." The other flat bar was describing observable behavior.

They are the same two criteria most managers miss.

  September 22, 2026

Key takeaways

Key takeaways
  • Observable behavior is what a camera would have recorded. Interpretations of motive or attitude are not observable, and they invite argument instead of change.
  • Impact is what that behavior cost the work, the team, or the customer. Without it, feedback reads as a preference the employee can safely deprioritize.
  • Agreement is the most misleading signal in a feedback conversation. People agree to end discomfort, not because they understood.
  • The failure is almost never knowledge. Managers can recite SBI and still skip both parts when a direct report gets emotional.
  • Structure becomes automatic through repeated scored practice against pushback, not through another workshop.

The agreement trap

Here is the conversation, compressed. A manager raises a missed report deadline with a direct report named Anya.

Anya blames a peer, Taylor. The manager acknowledges the frustration, asks about the dynamic, and proposes a three-way meeting.

Anya resists, worried it will become a blame session. The manager commits to keeping it forward-looking.

Anya agrees to a meeting time.

Read that back and try to find the moment where Anya learned what she did wrong. It is not there.

She was never told which specific action was the problem, and she was never told what it cost. What she was told, implicitly, is that a meeting will fix it.

The agreement trap is the tendency to read a cooperative ending as a successful conversation. Agreement measures how comfortable the exchange was, not whether the other person left knowing which behavior to change and why it mattered.

This is why manager self-assessment runs high and manager scores run low. The felt experience of a conversation that ends well is indistinguishable from the felt experience of a conversation that worked.

Only a rubric separates them, which is the broader argument in what your LMS can't show you about soft skills.

The scorecard in question ran six criteria. Four came back strong: opening and context setting, inviting dialogue and active listening, collaborative next steps, and a partial credit on describing the situation.

The manager was warm, asked a good question, and closed with a confirmed meeting. That profile is exactly the one that never gets corrected, because everything visible about it looks like competence.

45%
score on a feedback conversation that ended in agreement and felt successful
2 of 6
rubric criteria left effectively unaddressed, both of them conversation habits
0
specific behaviors named across the entire exchange
Rapport AI roleplay learner scorecard showing an overall score of 45 percent with no credit for describing observable behavior or explaining the impact
The real scorecard from the ATD Demo Day roleplay: strong marks for opening, dialogue, and next steps, and empty bars for observable behavior and impact.

Observable behavior, precisely

Most managers believe they describe behavior. What they usually describe is a conclusion drawn from behavior, and the difference decides whether the conversation becomes collaborative or adversarial.

The test is simple. Could a camera have recorded it?

If the statement involves a motive, an attitude, a pattern, or a character trait, the camera could not, which means it is an interpretation and the employee is entitled to dispute it.

Interpretation (invites argument)
Observable behavior (invites change)
"You are not taking deadlines seriously."
"The Q3 report was due Monday morning and arrived Thursday afternoon."
"You are being difficult with Taylor."
"In the last two project standups, you responded to Taylor's status updates by saying the timeline was unworkable, without proposing an alternative."
"Your communication needs work."
"The handoff email on the 14th listed the deliverables but not the owner or the date, so two people started the same task."
"You seem checked out lately."
"You have joined the last three client calls with video off and have not spoken unless asked a direct question."
Figure 1. Everything in the right column is a fact the employee can verify or correct. Everything in the left column is a verdict they can only accept or fight.
A manager with a notebook giving feedback to a direct report in a bright office meeting area
Specific beats general. The right column in Figure 1 takes more effort to say, and it is the version that changes behavior.

Notice how much more work the right column is. It requires the manager to have actually noticed something specific, remembered when it happened, and be willing to say it plainly.

That effort is precisely why the left column is so common under time pressure.

Why impact gets dropped, and what it costs

Impact is the second half of the sentence, and it is the half that answers "so what." It states what the behavior did to the work, the team, the customer, or the manager's ability to plan.

Managers drop it for an understandable reason: it feels like escalation. Naming a consequence sounds like an accusation, especially when the other person is already defensive.

So the manager softens, pivots to solutions, and moves on. That is what happened in the demo conversation.

The moment Anya got frustrated, the manager jumped straight to proposing a meeting.

The cost is that feedback without impact is indistinguishable from a preference. Compare:

  • Behavior only. "The report came in Thursday instead of Monday." The employee hears a scheduling observation.
  • Behavior plus impact. "The report came in Thursday instead of Monday, so the client review was pushed a week and two other people reworked their sections to match the new timeline." The employee hears a reason.

The second version is not harsher. It is more specific, and specificity is what makes a change feel necessary rather than optional.

Managers who want to rehearse exactly this moment can run the scenario in downward feedback or performance review.

Try it before your next one-on-one

Hold a feedback conversation with an avatar who deflects, blames a peer, and negotiates. Get scored on behavior and impact in about five minutes.

Run a difficult conversation

The same conversation, rewritten

Nothing below requires information the manager did not already have. It is the same facts, sequenced by someone who has had the conversation before.

What the manager said
What scores well
"I wanted to talk about that report delivery. What happened there?"
"I want to talk about the Q3 report. It was due Monday and came in Thursday afternoon. Walk me through the week."
"Yeah, that's really frustrating, I can tell."
"That sounds frustrating. Before we get to Taylor, I want to name the impact: the client review moved a week, and Priya and Marcus both reworked their sections."
"What exactly is the dynamic between you and Taylor?"
Same question, kept. This one was good, and it is the only line that made the direct report explain rather than defend.
"How about the three of us meet and I can try to facilitate?"
"Here is what I need regardless of how the Taylor conversation goes: if a deadline is going to slip, I need to hear it before the due date, not after. Can you commit to that?"
"How about we meet Wednesday at 2? How's that work?"
Same close, kept, but now it resolves a named behavior with a stated impact rather than standing in for one.
Figure 2. The rewrite is barely longer. It adds one behavior statement and one impact statement, which is the entire difference between a 45 and a passing score.
SBI feedback model diagram: situation, the Q3 report due Monday; behavior, arrived Thursday afternoon; impact, client review moved a week and two people reworked their sections
The SBI feedback model applied to the missed deadline. Behavior and impact are the two bands managers skip most.

Also worth noticing: the rewritten version still ends with the same scheduled meeting. Facilitating the peer conflict was a reasonable instinct.

It was just deployed as a substitute for feedback rather than in addition to it.

Building the habit, not the knowledge

Nothing in this article is new to most managers. SBI is taught in nearly every management curriculum, and the executive who scored 45% could almost certainly have explained the framework flawlessly ten minutes before running the scenario.

That is the point. The constraint is not recall in a calm room.

It is recall when a direct report gets emotional and starts assigning blame, at which point the prepared structure competes with the urge to smooth things over. Smoothing usually wins.

Three things make the structure automatic:

  • Pushback that feels real. Practice against a persona that deflects, holds a position, and concedes only on terms. A cooperative practice partner trains nothing, which is the distinction drawn in AI roleplay vs manager roleplay.
  • Scoring against the rubric you actually use. A general impression of how it went is what created the problem. Per-criterion scoring names the miss, and teams can score against SBI or upload their own framework.
  • Enough repetitions to matter. One attempt produces a score. Several attempts produce a habit, which is why the loop needs to be short enough to run again immediately. The science of engaging soft skills training covers the retrieval practice research behind this.

For L&D teams, the useful byproduct is visibility. When every manager in a cohort runs the same scenario, the dashboard shows whether impact statements are an individual gap or a program-wide one.

We cover reading those patterns in completion rates are not readiness and skill gap analysis.

A manager wearing headphones practicing a feedback conversation at his desk after hours
Private, repeated practice is what makes the structure hold when a real conversation gets tense.

Frequently asked questions

What is observable behavior in feedback?

A specific action a camera would have captured: what the person did, said, or delivered, and when. It is distinct from interpretation, which describes motive or character.

"The draft arrived Thursday instead of Monday" is observable. "You are not taking deadlines seriously" is not.

What is the SBI feedback model?

Situation, Behavior, Impact. Name the specific situation and when it happened, describe the observable behavior without interpretation, then explain the concrete impact.

The structure keeps feedback factual rather than evaluative, which lowers defensiveness and makes the requested change clear.

Why do managers skip the impact statement?

Because it feels confrontational, particularly when the employee is already upset. The cost is that feedback without impact reads as a preference rather than a requirement, and preferences get deprioritized.

Does agreement mean a feedback conversation went well?

No. Agreement often measures how quickly the discomfort ended.

An employee can agree to a next step without knowing which behavior triggered the conversation or why it mattered.

How do you practice giving feedback without practicing on your team?

Simulation. AI roleplay puts the manager in the conversation with a reactive avatar that deflects and negotiates like a real direct report, scores the attempt against a rubric such as SBI, and lets the manager rerun it immediately with the correction applied.

How many feedback conversations does a manager need to practice?

Enough that the structure survives emotional pressure. Framework knowledge is rarely the gap.

Recall under load is, and that is built through short, repeated, scored attempts rather than a single workshop.

Where to go next

The session this article is drawn from is recapped in full in our ATD Demo Day recap, including the live scenario build and the cohort dashboard. For the scenario design side, see how to design an AI roleplay scenario.

For the format fundamentals, start with AI role play training and communication role play scenarios that change behavior. L&D leaders building the case internally will want the business cost of skipping soft skills and the soft skills training overview.

Your managers already know the framework. Find out if they use it.

Run a scored feedback conversation against an avatar who pushes back, and see which criteria hold up when the room gets tense.

Try a feedback scenario For L&D teams