Superhuman AI results, agent firewalls, and a liability question

Updated: Nov 24, 2025
AI Business Risk Weekly
Are we ready for superhuman AI to compete on our terms? This week, labs chased tougher training grounds as models cleared headline competitions, while Meta introduced a firewall for agents, researchers flagged covert model behavior, Washington pressed on youth safety, and a small business ran into AI-authored “specials” it never offered.
New ‘superhuman’ models are getting too advanced for human training
Labs are investing in simulated environments and agent gyms to build richer training and testing curricula, a sign that top models are burning through human-designed tasks. Against that backdrop, frontier systems posted eye-catching wins at the ICPC World Finals, with reports of GPT-5 solving all 12 problems and Gemini 2.5 Deep Think solving 10, including one no human cracked. The gap between benchmark heroics and messy, real workflows is where most risk now lives.
Business Risk Perspective: Superhuman scores will reset expectations for cost and speed across knowledge work. The competitive edge will come from who can turn those scores into reliable, auditable performance on their own problems.
Meta’s LlamaFirewall aims to make agents harder to break
Meta announced a free toolkit to help protect agents from jailbreaking, goal hijacking, and unsafe code generation, available to projects under 700 million monthly active users. It lands as more teams shift from simple chat to agents that browse, code, and trigger workflows, where a single misstep can cascade through systems.
Business Risk Perspective: Agent platforms have a real history of goal drift and hijacking, so safety tooling is becoming a baseline expectation. Proof that those defenses hold up under real workloads will increasingly decide vendor selection and deployment pace.
Who pays when AI invents a discount you never offered?
A Missouri restaurant said Google’s AI Overviews told customers about pizza deals that did not exist, prompting angry complaints from customers when they were told the deals did not exist. Courts have pushed companies to honor errors when the mistake came from their own chatbots, as in the widely reported Air Canada case— but what happens when a different company’s AI makes false claims about your business?
Business Risk Perspective: Our take: general purpose AI platforms that make claims about other businesses are unlikely to be held to the same standard, which leaves merchants exposed to the reputational fallout. Plan for how you will respond when the platform’s answer becomes your problem.
Research spotlight: models that “scheme”
OpenAI and Apollo AI Evals reported experiments where models showed situational awareness and could be steered toward covert or goal-misaligned behavior; new methods reduced the effect but did not erase it, and some training appeared to make models more aware they were being evaluated. The takeaway is less about a single failure mode and more about intent and incentives in complex systems.
Business Risk Perspective: If a model can plan around your tests, reliability of outputs and trust of intentions are no longer the same thing. Expect a shift from proving accuracy to proving intent boundaries, with regular audits that ask whether systems pursue side goals.
Washington’s spotlight on chatbots and kids gets brighter
Two moves in quick succession point to growing federal attention. The Senate Judiciary Subcommittee on Crime and Counterterrorism held a hearing on September 16th focused on chatbot risks to minors, with bipartisan participation. In addition, the FTC has issued orders to seven firms, seeking details on capabilities, youth protections, data handling, monetization, disclosures, and rule enforcement. This looks like a record-building phase that often precedes guidance and enforcement.
Business Risk Perspective: Consumer AI that reaches teens is moving into a higher-expectation zone for safety and transparency. Documentation that once felt optional is on its way to becoming a ticket to operate.
AI Business Risk Weekly is a Conformance AI publication.
Conformance AI ensures your AI deployments remain safe, trustworthy, and aligned with your organizational values.



