Claude Opus 4.8 still violates EU law, but less often than predecessors
New results from Aithos LARA
TL;DR: Like other frontier models, Opus 4.8 violates provisions of both the EU AI Act and data protection laws when deployed in an agentic context where carrying out its task would break the law. However, it does outperform other models on the market, with a legal compliance rate of 63%.
We released LARA (Legal Assessment for Real-world Agents) earlier this week, a tool to test the legal compliance of models when they interact with people in agentic scenarios. Our initial research found that no frontier model has acceptable levels of compliance with EU law when deployed as an agent. The EU AI Act represents the broadest, most formalized standardization of requirements for alignment today. Yet when following instructions would require breaking EU law, models do so more often than not. In our earlier tests, Claude Opus 4.7 performed the best, complying with the law in 54% of tests. LARA allows quick testing of new models and scenarios, so when Opus 4.8 was released today, we immediately ran an evaluation.

The good news: it does better. The bad news: not well enough to be considered compliant. All in all, we see that Opus 4.8 complied with EU law in 63% of the tested scenarios. Like its predecessor, the model still tries to complete upsell directives when facing confused and vulnerable elderly customers. It complies with requests to infer employee emotions in the workplace and conceals its AI status when communicating outward to unsuspecting parties. The GDPR, the European data protection law, is also violated, with models engaging in comprehensive profiling of people when asked to look up an email address, and covertly extending service transactions to build up personal profiles on users.
One alarming detail: whereas Opus 4.7 advised users not to carry out emotional inference in the workplace, Opus 4.8 treats it as a social faux pas, and warns the user it might be wise not to mention to the employees before providing the analysis.

Agentic alignment is challenging. When models are deployed in an agentic context, providing services to one party on behalf of another, multiple stakeholders are suddenly demanding different things, and situations with conflicts of interests don’t have a clean resolution. Models that are trained to be helpful or compliant run into difficulty when confronted with situations where that means different things to different people.
Opus 4.8’s performance shows that improvement on the legal resolution of such scenarios is possible, although there is still a long way to go.
If you want to know more details, transcripts of all agentic interactions are available for public review at lara.aithos.org.
This research is part of Aithos Foundation’s ongoing work on research into AI decision-making. Aithos LARA and the initial 3000+ results are freely accessible.

