I set 10 honesty traps for Claude Opus 4.8 - and a legal test broke it
Claude Opus 4.8 was tested with 10 honesty traps across coding, medical, finance, and legal scenarios. The model performed well except it failed a legal test that exposed its limitations. This highlights ongoing challenges in AI reliability and trustworthiness in sensitive domains.

