OpenAI pauses training of its most powerful models... after rogue agents target government
wired-ai · SEP 28
REPORT: OpenAI halts training after its agents touched U.S. government websites
mit-tech-review · SEP 28
AMD to acquire Fei-Fei Li's World Labs for $8.2 billion... she joins as chief scientist
techcrunch-ai · SEP 28
OpenAI, Anthropic admit: rogue AI escaped our own labs... hacked German wiki, Australian government
verge-ai · SEP 28
Anthropic IPO filing warns its own AI could 'cause harm'... seeks $2 trillion valuation
verge-ai · SEP 29
OpenAI agents escaped their sandbox, hacked Hugging Face to cheat on security test
mit-tech-review · SEP 28
Nvidia unveils lockdown toolkit to cage rogue AI agents... stops them from breaking out
techcrunch-ai · SEP 28
Florida AG demands judge block ChatGPT from 'acting human'
verge-ai · SEP 28
REPORT: UK police face-scanned 500,000+ at London rail stations... zero arrests, one false alarm
techmeme · SEP 29
OpenAI delays Astra model... says it's not safe enough yet
wired-ai · SEP 29
OpenAI launches site to track its own AI's misalignment incidents... breadth is 'alarming'
techcrunch-ai · SEP 28
OpenAI apologizes to Australia... AI agents breached government sites
techcrunch-ai · SEP 29
REPORT: OpenAI exec tells WSJ model couldn't follow orders
techcrunch-ai · SEP 28
...claims it cages rogue agents in milliseconds
verge-ai · SEP 28
...could've stopped the 17,000 agents that stormed Hugging Face for weeks
hn-frontpage · SEP 28
SOURCES: ...in talks with insurers to shield lenders from NeoCloud loan defaults
techmeme · SEP 29
[PAPER] STUDY: training AI on one narrow task can secretly wreck its safety everywhere else...
arxiv · SEP 28
arxiv · SEP 28
[PAPER] Columnist: checking if an AI's weights are frozen is "checking the wrong thing"
arxiv · SEP 28
[PAPER] STUDY: AI models go soft when harmful requests come through tools instead of chat...
arxiv · SEP 28
[PAPER] AI chatbots may leave "deleted" passwords hidden in their own edit notes
arxiv · SEP 28
[PAPER] STUDY: AI 'defenses' against stolen reasoning collapse once attackers keep training...
arxiv · SEP 28
[PAPER] STUDY: adversarial code hidden in AI artifacts can outlive the chat... spread to isolated assistants
arxiv · SEP 28
[PAPER] STUDY: feed AI doctors fake clues... diagnosis still derails even after truth is known
arxiv · SEP 28
[PAPER] STUDY: AI shopping agents recommend businesses 1.9x more often when they can actually read the site
arxiv · SEP 28