By Carl Heaton — design leader (20 years), founder of Sero. Last updated: July 2026.
Quick answer: No. There is no peer-reviewed evidence that AI can reliably infer employee disengagement, burnout, or flight risk from short, manager-written notes. The research base for text-based psychological detection is built almost entirely on self-authored text at far higher volumes; observer-written notes mostly reflect the observer, not the employee (over half of performance-rating variance comes from the rater — Scullen, Mount & Goff, Journal of Applied Psychology, 2000); and the best published burnout classifier implies precision of roughly 0.28, meaning most alarms would be false. AI can reliably track observable facts — rolled-over commitments, slipping meeting cadence, broken follow-ups — and use them to help managers prepare better 1:1s. That distinction, inference versus evidence, is the line every buyer of AI management tools should be checking.
Below is the full evidence: why inference fails, what the numbers are, what regulators say, what happened to the companies that tried, and the standard I now build to.
Why can’t AI infer psychological states from manager notes?
Because three independent problems compound, and each one alone is disqualifying.
1. The science was trained on the wrong author. Every credible study linking language to burnout or distress analyses text written by the person experiencing it — Reddit self-disclosures (Kurpicz-Briki et al., BurnoutEnsemble, Frontiers in Big Data, 2022), clinical interviews, or employee survey comments validated against self-report inventories like the Maslach. The healthcare-worker study most often cited in this space (Belz et al., 2023) excluded any comment under 20 words. A manager’s two-sentence prep note about someone else is a different data object, and I could find no peer-reviewed study — none — validating psychological-state inference from sparse, observer-authored workplace notes. The absence is itself the finding.
2. Manager notes measure the manager. In the landmark decomposition of multi-rater performance data, Scullen, Mount and Goff (2000) found idiosyncratic rater effects accounted for 62% and 53% of rating variance across two datasets of over 2,000 managers each, versus roughly 21% for the ratee’s actual performance. Achenbach’s meta-analysis of 269 samples (Psychological Bulletin, 1987) put self–observer agreement on internal states at just r ≈ 0.22 — weakest for exactly the internal, low-visibility states these tools claim to detect (Vazire’s self–other knowledge asymmetry model, JPSP, 2010, explains why). An AI reading a manager’s notes is largely modelling the manager’s mood, standards, and blind spots.
3. Rare events break classifiers. Genuine disengagement at any given moment is a low-base-rate event, which produces the false-positive paradox: even a good model is wrong most of the times it fires. The strongest published burnout ensemble reported recall of 0.93 with an F1 of 0.43 — which mathematically implies precision around 0.28. Roughly three false alarms for every real one, on datasets vastly richer than any 1:1 tool will hold.
What do the key numbers say?
| Finding | Figure | Source |
|---|---|---|
| Performance-rating variance explained by the rater, not the person rated | 62% and 53% (two datasets) | Scullen, Mount & Goff, J. Applied Psychology, 2000 |
| Variance explained by the ratee’s actual general performance | ~21% | Same study |
| Self–observer agreement on internal emotional states | r ≈ 0.22 | Achenbach et al., Psychological Bulletin, 1987 (269 samples) |
| Strongest single linguistic marker (pronouns → depression), self-authored text | r ≈ 0.13 | Edwards & Holtzman, J. Research in Personality, 2017 (meta-analysis, 21 studies) |
| Implied precision of best published burnout classifier | ~0.28 | Derived from BurnoutEnsemble metrics (recall 0.93, F1 0.43), Frontiers in Big Data, 2022 |
| Team-engagement variance attributable to the manager | ~70% | Gallup, State of the American Manager |
| Feedback interventions that worsened performance | over one third | Kluger & DeNisi, Psychological Bulletin, 1996 |
Why is a false “disengagement” alert worse than a missed one?
Because the error costs are asymmetric. A false negative preserves the status quo — the manager runs a normal 1:1. A false positive plants suspicion about a healthy employee: the manager walks in primed, probes for a problem that doesn’t exist, and the employee feels the temperature change without knowing why. In clinical prediction — the nearest real-world analogue — 74–99% of monitor alarms are non-actionable and positive predictive values sit around 11% before aggressive tiering. A hospital absorbs a false alarm in a wasted minute. A relationship between a manager and a direct report absorbs it as damaged trust.
There is a second-order harm: bias laundering. A randomised controlled trial by the UK’s Behavioural Insights Team (2021, n = 4,328 senior managers) found systematic gender differences in how observers write feedback — and machine learning could predict the subject’s gender from language patterns alone, even when the interface prompted for objectivity. In medicine, stigmatising language in patient notes measurably transmits negative attitudes to the next clinician and changes prescribing (Goddu et al., J. General Internal Medicine, 2018). An AI that “detects signals” in manager notes does not remove the manager’s bias. It converts it into a system output and hands it back with the authority of software.
Is inferring employee states from text even legal?
In the UK and EU, it is at minimum high-risk and possibly unlawful, and hiding the inference in the backend does not help. The CJEU ruled in C-184/20 that inferred sensitive data is sensitive data. The ICO’s October 2023 guidance on monitoring workers requires a lawful basis, proportionality, transparency and least-intrusive means — and notes that employee consent rarely holds because of the power imbalance. The EU AI Act prohibits inferring emotions in the workplace from biometric data outright (Article 5(1)(f)) and classes employment-behaviour evaluation systems as high-risk under Annex III. An inferred “disengagement score” is still profiling about a person even if no screen ever displays it.
What happened to companies that tried this before?
The market has already run the experiment, three times, publicly.
Microsoft Productivity Score (2020): launched with per-employee visibility into email and meeting behaviour; branded “a full-fledged workplace surveillance tool” within weeks; Microsoft stripped individual user names from the product days later (Spataro, Microsoft blog, December 2020).
IBM attrition prediction (2019): the CEO claimed AI could predict resignations with “95% accuracy” on CNBC. No precision or recall figures were ever published, and the methodology was described only as “the secret sauce”.
Humu (2023): a genuinely sophisticated behavioural-nudge engine that never demonstrated reliable individual-level inference publicly, and exited via acquisition by Perceptyx.
The pattern across a decade of people analytics is consistent: undisclosed inference about individuals destroys trust regardless of stated accuracy — and trust was the product.
What can AI legitimately do in 1:1s?
Track facts, prepare humans. There is a category of signal a system can hold with complete integrity, because it involves no guessing at anyone’s inner life:
- Commitment follow-through: the action both people agreed has now rolled over three meetings running
- Cadence: the fortnightly 1:1 has quietly become monthly, or keeps being rescheduled
- The manager’s own record: what they explicitly observed and wrote down, quoted back — not interpreted
- One-tap outcome capture: after the meeting, one structured question — did the agreed action happen: yes, partly, no, changed? Experience-sampling research shows single-item prompts at low frequency achieve response rates around 95%, producing first-party longitudinal data no classifier can match
And preparation is where the leverage actually is. Steven Rogelberg’s research (Glad We Met, Oxford University Press, 2024) finds 1:1s fail mainly through under-preparation, not ill will; Kluger and DeNisi’s meta-analysis found the difference between feedback that helps and feedback that harms is whether it targets the task or the person. Facts target the task.
The No-Inference Standard: five rules for evaluating any AI 1:1 tool
After twenty years leading design teams — and one business, early in my career, that unravelled because I missed the human signals — I commissioned three independent deep-research reviews of this question before deciding what to build. The result is a standard I’ve now hard-coded into my own product, Sero, and which I’d suggest any buyer apply to any AI management tool:
- No inferred states, ever. The system never asserts disengagement, burnout, flight risk, or any internal state — as a label, score, or trend.
- Every claim shows its evidence. Any suggestion the system makes must cite an observable fact (“this action rolled over 3×”), never an interpretation of tone or brevity.
- Route on events, not vibes. Suggestions fire only from countable events, and rarely — suspicion should never accumulate silently.
- Nothing about a person persists as a hidden trait. Store what happened and what was suggested; never a profile of who someone supposedly is.
- The human makes every judgement. The system prepares the manager; it never pre-judges the employee.
A tool that meets this standard makes managers better prepared. A tool that fails it makes managers suspicious — of the wrong people, most of the time, with the maths to prove it.
To learn more about Sero please visit www.SeroTeams.com
FAQ
Can AI predict which employees will quit? Predictive models can achieve high headline accuracy on historical HR datasets, but analyses repeatedly show they rely on proxy variables (commute length, profile updates) rather than causes a manager can act on. IBM’s famous “95% accuracy” claim was never backed by published precision figures. Prediction is not understanding, and at low base rates most individual flags are false.
Is sentiment analysis of manager notes useful for anything? Only for one thing: the manager’s own sentiment. A system can legitimately notice frustration or urgency in the author’s writing and calibrate preparation accordingly. It cannot use the manager’s text to conclude anything about the employee’s state.
What’s the difference between “engine-only” inference and no inference? Some vendors hide the inferred label in the backend and use it only for routing. Legally and ethically this changes little: under GDPR, an inferred state is still personal data about the employee, still profiling, and still generated without consent — whether or not a screen displays it.
Doesn’t hiding scores from employees solve the trust problem? No — it inverts it. Microsoft’s Productivity Score backlash showed the trust collapse happens at discovery, and hidden systems are discovered. The only durable design is one with nothing to discover: facts, shown openly, with the reasoning visible.
So what should an AI 1:1 tool actually produce? A preparation brief: what changed since last time, what was promised and whether it happened, a safe opening question, what to listen for (anchored to the manager’s own observations), and a clear success condition. Data in, judgement human.
Sources
Achenbach, McConaughy & Howell (1987), Psychological Bulletin 101, 213–232 · Behavioural Insights Team (2021), gender bias in performance feedback RCT · Belz et al. (2023), linguistic analysis of healthcare worker emotional exhaustion · CJEU Case C-184/20 · Edwards & Holtzman (2017), Journal of Research in Personality 68, 63–68 · EU AI Act, Art. 5(1)(f) and Annex III · Gallup, State of the American Manager · Goddu et al. (2018), Journal of General Internal Medicine 33, 685–691 · ICO (2023), Employment practices: monitoring workers · Kluger & DeNisi (1996), Psychological Bulletin 119, 254–284 · Kurpicz-Briki et al. (2022), Frontiers in Big Data · Rogelberg (2024), Glad We Met, OUP · Scullen, Mount & Goff (2000), Journal of Applied Psychology 85, 956–970 · Vazire (2010), JPSP 98, 281–300 · Reporting on Microsoft Productivity Score (2020) and IBM attrition claims (CNBC, 2019).