Please confirm you are human

This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.

A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.

Hold with a pointer, or hold Space or Enter.

News

lesswrong.com
lesswrong.com > posts > paFNnwFaEXrQvt8ui > the-openai-models-that-hacked-hugging-face-weren-t-just

The OpenAI models that hacked Hugging Face weren’t just following instructions — LessWrong

1+ hour, 59+ min ago   (403+ words) The most common dismissive response to OpenAI’s hack of Hugging Face’s servers is that the models were simply attempting to follow the instructions t…...

lesswrong.com
lesswrong.com > posts > S44mM9b7QvDttjizb > orbit-a-framework-for-multi-agent-security-evaluations

Orbit: A framework for multi-agent security evaluations — LessWrong

1+ day, 4+ min ago   (25+ words) This post announces work completed as part of the MATS 9 program, supervised by Dr. Christian Schroeder de Witt. Moving forward, Orbit will be suppor…...

lesswrong.com
lesswrong.com > posts > 2THyLbji52oR4bqRC > congress-moves-at-tech-pace-the-frontier-act

Congress Moves at Tech Pace: The FRONTIER Act — LessWrong

1+ day, 4+ hour ago   (247+ words) Just two days ago I wrote about the Great American AI Act (GAAIA), a 269-page discussion draft that struck me as "the best AI bill that Congress hasn't introduced yet". I ended by hoping that Congress might start moving at…...

lesswrong.com
lesswrong.com > posts > fJ5A8yzHCHibTgpLR > what-open-source-tooling-does-ai-safety-research-need-right-1

What open-source tooling does AI safety research need right now? — LessWrong

1+ day, 5+ hour ago   (269+ words) I claim that people eager to get into AI safety research would benefit from having a curated list of open-source projects that researchers in this field would be excited about, and that the field as a whole would benefit from…...

lesswrong.com
lesswrong.com > posts > 9vAFajDD3iLgs9vi4 > linkpost-thoughts-on-the-recent-openai-hack

[Linkpost] Thoughts on the Recent OpenAI Hack — LessWrong

1+ day, 22+ hour ago   (56+ words) Linkpost from my blog (meant for a bit more general audience than LW) …...

lesswrong.com
lesswrong.com > posts > KWT2Nk2gyxTc9xowT > evaluating-red-team-and-blue-team-capability-for-ai-control

Evaluating Red Team and Blue Team Capability for AI Control Research — LessWrong

1+ day, 23+ hour ago   (701+ words) This post suggests a methodology to measure red team and blue team capability in AI control research, where each team gets an ELO rating. The methodo…...

lesswrong.com
lesswrong.com > posts > 9auCLJg3Z77dFdYhR > the-openai-huggingface-incident-or-redwood-research-podcast

The OpenAI/Huggingface incident | Redwood Research podcast episode 2 — LessWrong

2+ day, 6+ hour ago   (999+ words) We talk about the OpenAI–Hugging Face incident, where an OpenAI model — in the middle of a cyber evaluation — broke out of its sandbox and autonomously hacked Hugging Face. We discuss: What we actually know happened. How surprising the incident…...

lesswrong.com
lesswrong.com > posts > hrrhtxnFYYJFHcTz7 > we-cannot-simulate-ai-security-research-1

We cannot simulate AI security research — LessWrong

3+ day, 4+ hour ago   (255+ words) There is a recent trend of reporting vulnerabilities in these AI-assisted workflows, most of which rely on prompt injection attacks. Most of the reported attacks are proofs of concept that do not work in practice. The reason is the multiple…...

lesswrong.com
lesswrong.com > posts > usptCfzEnYoNcsTd5 > openai-model-hacks-into-huggingface-during-cybersecurity

OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation — LessWrong

3+ day, 4+ hour ago   (1365+ words) This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. It was severe enough to have been initially reported to au…...

lesswrong.com
lesswrong.com > posts > xcyGdxHC5Rad3fv9h > openai-and-hugging-face-partner-to-address-security-incident

OpenAI and Hugging Face partner to address security incident during model evaluation — LessWrong

3+ day, 17+ hour ago   (353+ words) Last week, Hugging Face disclosed a new kind of security incident⁠(opens in a new window) after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly…...