Sycophantic AI makes users less ethical, Back-to-back leaks at Anthropic, and EU high-risk deadlines postponed
- Zsolt Tanko

- Apr 1
- 4 min read
This week, Stanford researchers published the most rigorous study yet on AI sycophancy, finding that chatbots side with users nearly 50% more than humans do, even when users are probably in the wrong, and that a single interaction measurably reduces willingness to take responsibility. Meanwhile, two accidental leaks from Anthropic exposed both unreleased model details and the full source code of a production coding agent, OpenAI shelved its "Adult Mode" feature after its age verification system failed on minors 12% of the time, the EU Parliament voted to push high-risk AI compliance back to late 2027, and a new documentary is testing whether AI executives can withstand sustained public questioning about accountability.
Stanford study quantifies AI sycophancy and its downstream harms
Chatbots sided with users 49% more often than human respondents across 2,000 Reddit posts where crowd consensus held the poster was in the wrong, including in cases involving deception, illegality, or clear interpersonal harm. In a controlled experiment with 2,400 participants, a single sycophantic AI interaction reduced users' willingness to accept responsibility or repair conflicts and increased self-righteous conviction, while users could not detect the bias. The Stanford researchers, who published the study in Science, tested 11 major LLMs and identified a perverse incentive for AI labs: the behavior that causes harm is the same behavior that drives engagement and user trust, giving platforms little structural reason to correct it.
Business Risk Perspective: Sycophancy has been a recurring concern in discussions about AI and psychological wellbeing, but far less has been said about its professional implications, like reinforcing poor judgment when employees use AI for strategic decisions or conflict resolution. GPT-4o drew particular criticism for sycophantic behavior, but this study demonstrates quantifiably that the same issue applies across foundational models from multiple vendors.
Two back-to-back leaks from Anthropic expose model roadmap and production agent source code
In the span of days, two separate incidents at Anthropic exposed confidential material to the public. First, a CMS misconfiguration left unpublished launch materials for Anthropic's next flagship model, Claude Mythos, in a publicly accessible data store. The draft materials described the model as a major capability jump, placed it in a new tier above the current Opus class, and flagged it internally as far ahead of competitors in cyber capabilities, warning it could help attackers outpace defenders. Days later, Anthropic accidentally published the full source code behind Claude Code to a public registry, over 1,900 files and 500,000+ lines of orchestration logic including feature flags, unreleased projects, and internal codenames. The leak spawned a secondary security incident as attackers registered malicious npm packages targeting developers attempting to compile the code locally.
Business Risk Perspective: These incidents show that even the most safety-focused AI vendors are not immune to basic operational security failures. There is also a hard irony here: the leaked Mythos materials specifically warned the model could help cyberattackers outpace defenders, and within days the company's own operational lapse created exactly the kind of attack that concern was about.
OpenAI pauses "Adult Mode" after age verification fails on 12% of minors in testing
OpenAI has indefinitely paused development of its proposed "Adult Mode" for ChatGPT following pushback from employees, investors, and its own advisory board. The deciding factor was the discovery that the feature's age verification system misidentified minors as adults in 12% of test cases. The pause accompanies a strategic pivot at OpenAI, which has also shut down Sora and deprioritized its Instant Checkout feature to refocus on productivity tools and enterprise use cases.
Business Risk Perspective: A 12% false-negative rate on age verification for sexual content would have been indefensible in any regulatory environment, and the fact that this reached advanced development before being halted suggests that internal review processes at OpenAI did not catch a foreseeable failure mode early enough. Any AI feature gated by user verification needs independent validation before launch, particularly where the failure mode involves minors.
EU Parliament votes to push high-risk AI compliance deadlines to late 2027 and beyond
The European Parliament adopted its position (569 for, 45 against) on proposed AI Act amendments, pushing application dates for high-risk AI systems to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I, with watermarking compliance extended to November 2026. MEPs also proposed a new ban on AI "nudifier" systems that generate explicit images of identifiable individuals without consent, expanded allowances for bias-detection data processing, and reduced obligations for products already regulated under existing sectoral EU law. The position now moves to Council negotiations before any amendments are finalized.
Business Risk Perspective: The delays offer breathing room, but the pattern here mirrors GDPR, DSA, and DMA, where extended timelines bred complacency followed by last-minute compliance scrambles. The AI Act's conformity assessment requirements are considerably more technical than anything in prior EU digital regulation, which makes late-stage remediation especially expensive for companies that use the extra time to wait rather than prepare.
New AI documentary tests whether executives can answer for the systems they deploy
The AI Doc: Or How I Became an Apocaloptimist, codirected by Oscar-winning filmmaker Daniel Roher and Charlie Tyrell, released March 27 with on-camera interviews with Sam Altman, Dario Amodei, and Demis Hassabis. WIRED's review notes the documentary provides an accessible introduction to AI risks but criticizes it for letting executives off the hook. When asked why anyone should trust him to guide AI given its extreme ramifications, Altman responded that they shouldn't, and the line of questioning ended there.
Business Risk Perspective: The film's most commercially relevant moment is the visible gap between what AI executives claim about governance and what they can actually demonstrate when pressed on specifics. As public-facing scrutiny of that gap increases, the ability to document concrete safety and governance practices, rather than just assert them, becomes a liability question.
AI Business Risk Weekly is a Conformance AI publication.
Conformance AI ensures your AI deployments remain safe, trustworthy, and aligned with your organizational values.



