EU AI Act deadline holds, Anthropic models regress, and frontier cyber capabilities leak
- Zsolt Tanko

- Apr 29
- 4 min read
AI Business Risk Weekly
This week, the EU AI Act's August 2nd deadline for high-risk AI systems survived a 12-hour negotiation session that was expected to push it back. Meanwhile, Anthropic is dealing with user backlash on two fronts: widespread complaints that Opus 4.7 is a downgrade from 4.6, and a month of silent bugs in Claude Code. Some users are jumping to new open-weight models to escape the instability, but the hallucination rates are unacceptable for enterprise. And both Anthropic and OpenAI now have models with serious offensive cyber capabilities, with both revealing containment vulnerabilities.
EU AI Act high-risk obligations were expected to be delayed, but weren't
The EU AI Act's Annex III high-risk obligations are set to take effect on August 2nd, 2026, 95 days from now. They apply to AI used in hiring, credit scoring, insurance pricing, education admissions, and other uses deemed to be high-risk. If you're a provider of these systems, you need risk management systems, conformity assessments, CE marking, and EU database registration before deployment. If you're a deployer, you need human oversight, operational monitoring, log retention for at least six months, and must inform individuals when AI is involved in decisions affecting them. Fines
run up to €15 million or 3% of global annual turnover.
Many organizations haven’t been preparing, and of those that were, many compliance teams had been planning around a delay. Both Parliament and Council had agreed to postpone this deadline to December 2027 as part of the broader Digital Omnibus, but the April 28 trilogue ran for 12 hours and ended without a deal. Talks are set to resume mid-May.
Business Risk Perspective: If your compliance program was built around an assumed delay, reset it now (we can help, with minimal hours from your team). A deal in May is highly plausible, but it will still need to go through an IMCO/LIBE committee vote and a European Parliament plenary vote. August 2nd is only 95 days out, and the compliance work is identical either way.
Anthropic's reliability problems: Opus 4.7 backlash and Claude Code bugs
Opus 4.7 arrived to significant user backlash, with power users reporting worse reasoning, more hallucinations, and reduced control, with multiple thinking levels replaced with just one ‘adaptive thinking’ mode. A viral open letter alleged the model fabricated an entire chapter while editing a manuscript. The official model card shows improved efficiency and strong benchmarks, but the volume of complaints suggests a gap between benchmarks and reality, as we’ve seen with other models.
Separately, Anthropic acknowledged three bugs that degraded Claude Code over a month, including an unannounced reasoning-effort downgrade and a caching regression. Everything was fixed by April 20th, but users had been reporting the same symptoms for weeks before Anthropic said anything.
Business Risk Perspective: Foundation models will keep changing, and your system's behavior will shift even when you don't touch anything. This is part of why it’s important for organizations to continuously run their own evals, or have an independent third party running them.
Open models are closing the gap on reasoning, but not hallucination rates
Some users are migrating to new open-weight models for lower costs and more independence from changing foundation models. Open models are closing the capabilities gap with the frontier, and for the first time, many users are saying that they are viable alternatives for coding.
In the past two weeks, Moonshot released Kimi K2.6, claiming open-source state-of-the-art on several coding benchmarks, with some Opus subscribers publicly switching. Alibaba previewed Qwen3.6-Max with strong agentic coding. DeepSeek released V4 at $0.14/$0.28 per million tokens versus Opus 4.7's $5/$25.
That said, we cannot recommend this move for enterprise. One source, Artificial Analysis's AI Omniscience evaluation, measures how often models fabricate information (lower is better). The chart speaks for itself.
Business Risk Perspective: Open models provide lower cost, version control, and deployment autonomy, but the hallucination rates aren’t viable for enterprise contexts. Preparing for regular model updates from frontier labs may seem onerous, but with capabilities advancing so quickly, it’s not really optional to choose a model and stick to it (even if an open-source model technically gives you the choice.)
Two frontier labs have dangerous cybersecurity models, and both are vulnerable
A Discord group reportedly gained access to Anthropic's Mythos cybersecurity model within days of its limited release, reportedly by guessing deployment URLs using credentials from the recent Mercor breach. Mythos was deemed too powerful for public release and had triggered White House emergency meetings.
OpenAI took a different approach, releasing GPT-5.5 to the public with special guardrails around cybersecurity capabilities. Before launch, the UK AI Security Institute found a universal jailbreak after six hours of red-teaming, and assessed it as the most cyber-capable public model available. OpenAI says the jailbreak was fixed before launch, but AISI wasn't given access to verify the final configuration.
Business Risk Perspective: AI-powered cyberattacks are becoming a true threat, with the unreleased Mythos finding “thousands of zero-day vulnerabilities, many of them critical,” in just a few weeks. If the news from this week is any indication, organizations should be updating security postures and incident response plans to account for the growing risks.
AI Business Risk Weekly is a Conformance AI publication.
Conformance AI ensures your AI deployments remain safe, trustworthy, and aligned with your organizational values.



