Why Managers Score Lower Than They Think: The Two Things Missing From Most Feedback Conversations
At ATD Demo Day, a seasoned executive ran a live conflict resolution roleplay in front of a few hundred L&D leaders. The conversation ended in agreement.
A meeting got scheduled. Everyone was polite.
The scorecard came back at 45%, and the report named the reason in three words: "Impact not explained." The other flat bar was describing observable behavior.
They are the same two criteria most managers miss.
September 22, 2026
Key takeaways
- Observable behavior is what a camera would have recorded. Interpretations of motive or attitude are not observable, and they invite argument instead of change.
- Impact is what that behavior cost the work, the team, or the customer. Without it, feedback reads as a preference the employee can safely deprioritize.
- Agreement is the most misleading signal in a feedback conversation. People agree to end discomfort, not because they understood.
- The failure is almost never knowledge. Managers can recite SBI and still skip both parts when a direct report gets emotional.
- Structure becomes automatic through repeated scored practice against pushback, not through another workshop.
The agreement trap
Here is the conversation, compressed. A manager raises a missed report deadline with a direct report named Anya.
Anya blames a peer, Taylor. The manager acknowledges the frustration, asks about the dynamic, and proposes a three-way meeting.
Anya resists, worried it will become a blame session. The manager commits to keeping it forward-looking.
Anya agrees to a meeting time.
Read that back and try to find the moment where Anya learned what she did wrong. It is not there.
She was never told which specific action was the problem, and she was never told what it cost. What she was told, implicitly, is that a meeting will fix it.
The agreement trap is the tendency to read a cooperative ending as a successful conversation. Agreement measures how comfortable the exchange was, not whether the other person left knowing which behavior to change and why it mattered.
This is why manager self-assessment runs high and manager scores run low. The felt experience of a conversation that ends well is indistinguishable from the felt experience of a conversation that worked.
Only a rubric separates them, which is the broader argument in what your LMS can't show you about soft skills.
The scorecard in question ran six criteria. Four came back strong: opening and context setting, inviting dialogue and active listening, collaborative next steps, and a partial credit on describing the situation.
The manager was warm, asked a good question, and closed with a confirmed meeting. That profile is exactly the one that never gets corrected, because everything visible about it looks like competence.

Observable behavior, precisely
Most managers believe they describe behavior. What they usually describe is a conclusion drawn from behavior, and the difference decides whether the conversation becomes collaborative or adversarial.
The test is simple. Could a camera have recorded it?
If the statement involves a motive, an attitude, a pattern, or a character trait, the camera could not, which means it is an interpretation and the employee is entitled to dispute it.

Notice how much more work the right column is. It requires the manager to have actually noticed something specific, remembered when it happened, and be willing to say it plainly.
That effort is precisely why the left column is so common under time pressure.
Why impact gets dropped, and what it costs
Impact is the second half of the sentence, and it is the half that answers "so what." It states what the behavior did to the work, the team, the customer, or the manager's ability to plan.
Managers drop it for an understandable reason: it feels like escalation. Naming a consequence sounds like an accusation, especially when the other person is already defensive.
So the manager softens, pivots to solutions, and moves on. That is what happened in the demo conversation.
The moment Anya got frustrated, the manager jumped straight to proposing a meeting.
The cost is that feedback without impact is indistinguishable from a preference. Compare:
- Behavior only. "The report came in Thursday instead of Monday." The employee hears a scheduling observation.
- Behavior plus impact. "The report came in Thursday instead of Monday, so the client review was pushed a week and two other people reworked their sections to match the new timeline." The employee hears a reason.
The second version is not harsher. It is more specific, and specificity is what makes a change feel necessary rather than optional.
Managers who want to rehearse exactly this moment can run the scenario in downward feedback or performance review.
Try it before your next one-on-one
Hold a feedback conversation with an avatar who deflects, blames a peer, and negotiates. Get scored on behavior and impact in about five minutes.
Run a difficult conversation →The same conversation, rewritten
Nothing below requires information the manager did not already have. It is the same facts, sequenced by someone who has had the conversation before.

Also worth noticing: the rewritten version still ends with the same scheduled meeting. Facilitating the peer conflict was a reasonable instinct.
It was just deployed as a substitute for feedback rather than in addition to it.
Building the habit, not the knowledge
Nothing in this article is new to most managers. SBI is taught in nearly every management curriculum, and the executive who scored 45% could almost certainly have explained the framework flawlessly ten minutes before running the scenario.
That is the point. The constraint is not recall in a calm room.
It is recall when a direct report gets emotional and starts assigning blame, at which point the prepared structure competes with the urge to smooth things over. Smoothing usually wins.
Three things make the structure automatic:
- Pushback that feels real. Practice against a persona that deflects, holds a position, and concedes only on terms. A cooperative practice partner trains nothing, which is the distinction drawn in AI roleplay vs manager roleplay.
- Scoring against the rubric you actually use. A general impression of how it went is what created the problem. Per-criterion scoring names the miss, and teams can score against SBI or upload their own framework.
- Enough repetitions to matter. One attempt produces a score. Several attempts produce a habit, which is why the loop needs to be short enough to run again immediately. The science of engaging soft skills training covers the retrieval practice research behind this.
For L&D teams, the useful byproduct is visibility. When every manager in a cohort runs the same scenario, the dashboard shows whether impact statements are an individual gap or a program-wide one.
We cover reading those patterns in completion rates are not readiness and skill gap analysis.

Frequently asked questions
What is observable behavior in feedback?
A specific action a camera would have captured: what the person did, said, or delivered, and when. It is distinct from interpretation, which describes motive or character.
"The draft arrived Thursday instead of Monday" is observable. "You are not taking deadlines seriously" is not.
What is the SBI feedback model?
Situation, Behavior, Impact. Name the specific situation and when it happened, describe the observable behavior without interpretation, then explain the concrete impact.
The structure keeps feedback factual rather than evaluative, which lowers defensiveness and makes the requested change clear.
Why do managers skip the impact statement?
Because it feels confrontational, particularly when the employee is already upset. The cost is that feedback without impact reads as a preference rather than a requirement, and preferences get deprioritized.
Does agreement mean a feedback conversation went well?
No. Agreement often measures how quickly the discomfort ended.
An employee can agree to a next step without knowing which behavior triggered the conversation or why it mattered.
How do you practice giving feedback without practicing on your team?
Simulation. AI roleplay puts the manager in the conversation with a reactive avatar that deflects and negotiates like a real direct report, scores the attempt against a rubric such as SBI, and lets the manager rerun it immediately with the correction applied.
How many feedback conversations does a manager need to practice?
Enough that the structure survives emotional pressure. Framework knowledge is rarely the gap.
Recall under load is, and that is built through short, repeated, scored attempts rather than a single workshop.
Where to go next
The session this article is drawn from is recapped in full in our ATD Demo Day recap, including the live scenario build and the cohort dashboard. For the scenario design side, see how to design an AI roleplay scenario.
For the format fundamentals, start with AI role play training and communication role play scenarios that change behavior. L&D leaders building the case internally will want the business cost of skipping soft skills and the soft skills training overview.
- We built an AI roleplay live at ATD Demo Day. The manager scored 45. (the full session recap)
- AI roleplay and employee data: 12 questions to ask before you roll it out
- How to design an AI roleplay scenario that actually changes behavior
- Completion rates are not readiness: what a cohort skill dashboard shows instead
Your managers already know the framework. Find out if they use it.
Run a scored feedback conversation against an avatar who pushes back, and see which criteria hold up when the room gets tense.
Try a feedback scenario → For L&D teams