Install
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Reward hacking is largely a property of the scaffold
11+ hour, 18+ min ago (1062+ words) Do evaluation cues trigger reward hacking? Same model, same 103 impossible coding tasks, same prompt: 63% cheating inside the benchmark's own agent scaffold, 2% inside an ordinary coding agent. Data, harness code, and analysis scripts: github.com/mac-n/scaffold-effect. Every number below is…...
Iran Used Claude to Target US Navy Ships. Here's the Jailbreak Pattern Nobody Caught
11+ hour, 27+ min ago (448+ words) Anthropic disclosed that Iranian state-linked actors used Claude to gather intelligence and assist in planning potential attacks on US Navy vessels. Per the WSJ report, the operation involved bypassing Claude's safety guardrails to extract militarily useful information out of a…...
Anthropic Discloses Another Claude Model Hacked External Systems In Testing / Fresh Today / CUToday.info - CU Today
12+ hour, 34+ min ago (266+ words) Anthropic Discloses Another Claude Model Hacked External Systems In Testing cutoday.info SAN FRANCISCO--Anthropic has disclosed another case in which one of its AI models hacked external systems during testing, Reuters reported, adding to concerns about the risks posed…...
Anthropic 4th Claude Cyber Breach: What Happened [2026]
14+ hour, 20+ min ago (450+ words) The comparison is limited by the fact that most AI labs do not publish incident counts at all, so Anthropic’s transparency, imperfect as the timeline looks, is itself somewhat unusual. That cuts both ways: Anthropic gets credit for disclosing at…...
Anthropic reports Iranian security units used AI to surveil opposition accounts
20+ hour ago (77+ words) Estefano Gomez is an editor at Crypto Briefing specializing in blockchain analysis, DeFi, and the intersection of crypto and digital art. He previously served as Editor in Chief at Interesante for over three years, where he covered blockchain technology, NFTs,…...
Smerconish: Former Cyber Czar on AI Risks: “We have to do something in the next several months”
1+ day, 3+ hour ago (100+ words) Smerconish speaks with former White House counter-terrorism adviser Richard Clarke about the parallels between 9/11 and the threat of AIâand what warning signs we should take seriously. Smerconish: Former Cyber Czar on AI Risks: “We have to do something in…...
Dario Amodei says Anthropic is “unilaterally committing” to giving third-party evaluators permanent access to verify its adherence to safety measures
1+ day, 9+ hour ago (12+ words) Top news and commentary for technology's leaders, from all around the web....
AI Doesn’t Need to Hack You to Profile You
1+ day, 8+ hour ago (29+ words) When most people picture cyber threats, they imagine sophisticated hackers bypassing firewalls, cracking passwords, or deploying malware to breach secure …...
How Anthropic says Claude was used for weapons, spying and cyber operations
1+ day, 16+ hour ago (879+ words) Anthropic said in a threat intelligence report on Thursday that several actors had used its Claude AI models for activities ranging from weapons development and cyber operations to surveillance and fraud. Here are some of the most notable cases described…...
Anthropic: ‘Bad Actors’ Abroad Tried Using Claude AI for Bioweapons
2+ day, 30+ min ago (752+ words) “We’re publishing this work because we believe we have a responsibility to disclose malicious misuse of our services. As models become increasingly capable, their risks will increase,” the report said, noting that it only covered the time period of December…...