Install
Safety Research
Interpretability, red teaming and the alignment work behind the models.
- 37 Tracked terms
- Last 30 days Feed window
What this topic collects on
An article joins this feed when it matches these terms. Each one is also a search of its own.
- anthropic
- claude
- dario amodei
- claude code
- model context protocol
- anthropics
- ai safety
- safety
- ai ethics
- ethics
- alignment
- red teaming
- guardrails
- research
- ai regulation
- ai policy
- ai-safety
- ai-ethics
- ai-regulation
- ai regulations
- ai-policy
- ai policies
- interpretability
- responsible scaling policy
- isbndb
- mrinank sharma
- destructive scanning
- recursive self-improvement
- ai training data
- ai oversight
- irregular
- rare books
- open-weight models
- u.s. ai competition
- constitutional ai
- mechanistic interpretability
- sleeper agents
Related topics
Latest in Safety Research
ChatGPT boss Sam Altman reveals the two ways AI progress 'could go very badly"
6+ hour, 33+ min ago (406+ words) ChatGPT boss Sam Altman has warned that the future of AI could go ‘very badly’ if the technology’s rapid progress continues to be developed. The OpenAI CEO shared his concerns in a series of posts on X, saying there were…...
Anthropic OpenAI Slow AI Development Push
2+ hour, 45+ min ago (893+ words) Companies like Anthropic and OpenAI face a complex challenge balancing calls to slow advanced AI development with pressure from tech sectors, financial markets, and President Donald Trump’s administration to maintain momentum. Anthropic and OpenAI are navigating a difficult equation between…...
Don???t spin us, guys
12+ hour, 28+ min ago (270+ words) Don’t spin us, guys The Times of India Here we go again Spinning however the AI overlords spin us They say the end of jobs is nigh Then they say, oh wait, no, the apocalypse is postponed They are supposed…...
New warnings about the risks of AI to humanity revive a long-running debate
4+ hour, 55+ min ago (877+ words) New warnings from the artificial-intelligence industry have revived debate over whether advanced AI models could escape human control and threaten humanity LOS ANGELES -- New warnings from within the artificial intelligence industry have revived a long-running debate over whether advanced AI…...
Ex-Anthropic researcher, who said 'AI could kill us all', calls for 'kill switch"
5+ hour, 3+ min ago (32+ words) Ex-Anthropic researcher, who said 'AI could kill us all', calls for 'kill switch' | 'Labs should regulate themselves' | Inshorts Inshorts Ex-Anthropic researcher, who said 'AI could kill us all', calls for 'kill switch'...
Former White House AI official urges OpenAI, Anthropic to slow AI development
9+ hour, 20+ min ago (255+ words) David Sacks, a former White House AI official, has lashed out at Dario Amodei (of Anthropic) and Sam Altman (of OpenAI), stating that they need to stop pretending they require permission to slow down AI development. Former White House AI…...
AI Leaders Urge Development Slowdown as Safety Warnings Mount
10+ hour, 18+ min ago (592+ words) streamlinefeed.co.ke A coordinated chorus of artificial intelligence executives and pioneering researchers is demanding an immediate slowdown in the training and release of frontier models, warning that runaway capability gains threaten to outpace human control. The warnings, punctuated by…...
Slow Down AI? Microsoft CEO, President Trump Line Up Against Anthropic's Call — Big Tech's Divide Is Now Out In The Open
9+ hour, 14+ min ago (716+ words) The AI industry is split over Anthropic’s call to slow the pace of frontier-model development, with OpenAI CEO Sam Altman and Microsoft CEO Satya Nadella backing more deliberate progress, while “The Big Short” investor Michael Burry and others question the…...
Amid AI Debate, Sam Altman Flags Two Ways Progress 'Could Go Very Badly"
9+ hour, 48+ min ago (666+ words) OpenAI CEO Sam Altman has flagged two ways in which the progress of artificial intelligence "could go very badly" amid the larger ongoing debate. Altman addressed both concerns in two posts on X as the debate over the pace of…...
"Gambling With Our Lives": Is AI on the Verge of Destroying Humanity?
9+ hour, 9+ min ago (989+ words) You may have noticed that major players in artificial intelligence have been tugging at their collars and giving each other nervous looks lately. Is it getting hot in here? Is AI on the verge of destroying humanity? The alarm bells…...