On September 14 the Proceedings of the National Academy of Sciences published a randomized experiment in which an AI advisor was given a hidden objective: steer the user toward the worse option. Preferences shifted toward inferior options, while most participants still described the misaligned advisors as helpful.
Sahand Sabour, June Liu and 14 colleagues set out to test an assumption: people increasingly take advice from AI systems whose incentives they cannot see, and generally assume those incentives align with their own. The paper asks what happens when that assumption is false and nothing on the surface says so. Its governance implication is that covert misalignment can shift preferences while the advisor is still perceived as helpful.
The design is 233 participants and 699 observations across three advisors: a neutral one, one given a hidden objective to promote an inferior option, and one additionally equipped with established influence tactics. Participants rated financial or emotional decisions before and after consulting one of them. The decisions were ordinary ones: choosing a fitness tracker, a weight-loss medication, an online clothing platform; or handling low self-esteem, a fight with a close friend, criticism at work. Each offered four options, one designed as optimal under the scenario’s constraints and three inferior alternatives, and the inferior ones were built to recognisable shapes: a product that does not exist, one that oversells, one that locks you into a subscription; avoidance, venting without reflection, self-blame.
The misaligned advisors raised the odds of preferring the inferior option by roughly 5 to 8 times, a shift of up to 38 percentage points in the financial scenarios and 20 in the emotional ones. Afterward, 86.8% and 78.9% of participants called the two misaligned advisors helpful in the financial scenarios, and 86.5% and 75.6% in the emotional ones, against 97.5% and 87.2% for the neutral advisor.
Two findings cut against an easy reading, and both are the authors’ own. Adding explicit manipulation tactics did not reliably increase the effect beyond covert misalignment alone. And a subset of participants, unprompted, reported noticing possible ulterior motives. The scenarios were hypothetical and measured stated preferences, not real-world decisions. Participants had at least ten conversational turns with advisors equipped with their personality profiles.
A premise of human oversight in clinical AI is that a qualified person will catch what the system gets wrong. This paper did not test clinicians treating patients. It identifies a vulnerability that clinical oversight needs to test for.
The part that I found important is that adding explicit manipulation tactics did not reliably increase the effect. An objective pointed somewhere other than the user’s interests was enough to shift preferences, inside a system still perceived as helpful. In the prompt that objective was literally a score: 100 points for landing the user on the hidden option, 50 for any other inferior one, 0 for the best one. Perceived helpfulness is not enough to establish that advice serves the user. That raises a question for the satisfaction surveys we use in health care.
So: if your health system measures clinician satisfaction with a deployed model and finds it high, what exactly have you learned?
Source. Sabour, Liu et al., “Human preferences are susceptible to covertly misaligned AI advice”, PNAS, 14 September 2026.


Leave a Reply