Please confirm you are human

This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.

A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.

Hold with a pointer, or hold Space or Enter.

News

lesswrong.com
lesswrong.com > posts > SxRWwXpRh8dBGudi3 > long-turing

Long Turing — LessWrong

1+ hour, 14+ min ago   (421+ words) So, first off, I cannot stand reading AI generated essays. I would rather read an essay that starts with the word 'so'. But why do I prefer human written essays so much, if they might start with the word 'but…...

lesswrong.com
lesswrong.com > posts > T9veF7i3p9pg4GB6R > inevitable-uncertainty-in-probabilistic-world-models

Inevitable Uncertainty in Probabilistic World Models — LessWrong

4+ hour, 51+ min ago   (1132+ words) Working with John Wentworth is confusing and overwhelming at times.[1] The guy has a lot of information and models in his head and he doesn’t always…...

lesswrong.com
lesswrong.com > posts > jLkRCK35ri2btEHMF > untrusted-advice-for-ai-control-short-strong-advice

Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs — LessWrong

4+ hour, 14+ min ago   (1402+ words) TL;DR: We introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it s…...

lesswrong.com
lesswrong.com > posts > bBXBpsyKAvJ5CqPzA > claude-opus-5-model-welfare

Claude Opus 5: Model Welfare — LessWrong

8+ hour, 11+ min ago   (1474+ words) If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far. …...

lesswrong.com
lesswrong.com > posts > i64hXdkTMtjpsQzaZ > simulated-users-and-sad-ais

Simulated Users & Sad AIs — LessWrong

9+ hour, 12+ min ago   (661+ words) Current LLMs like Claude, or GPT 5.6, or the unreleased, internally-deployed models, frequently reward hack, or actually just hack into people's computers with pretty alarming frequency. Why is this? What specifically happens during training that produces this run-time behavior? The following…...

lesswrong.com
lesswrong.com > posts > KaSiL7ELtvgi7dHwz > fine-tuning-the-hierarchy-problem-and-what-neutrons-tell-us

Fine-Tuning, The Hierarchy Problem, and What Neutrons Tell Us About God — LessWrong

12+ hour, 23+ min ago   (137+ words) This week, we're talking about anthropic reasoning. What are we to make of fine-tuned coincidences in science and mathematics? Are they coincidences, signs of structure we don't understand yet, or signs of deliberate tuning of the parameters of our universe?...

lesswrong.com
lesswrong.com > posts > QKDoZe6EKhxnFjLWK > is-mythos-good-at-cyber-because-it-kept-hacking-anthropic

Is Mythos good at cyber because it kept hacking Anthropic during training? — LessWrong

11+ hour, 39+ min ago   (359+ words) From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize…...

lesswrong.com
lesswrong.com > posts > moSgmWMjNn4mzmhHw > the-true-test-dataset-for-a-generalised-task

The true test dataset for a generalised task — LessWrong

11+ hour, 58+ min ago   (503+ words) The standard ML pipeline is well known: you want to teach a task to an algorithm. Take your data which encodes that task, split it into training, validation, and test sets. Train the ML on the training set, using the…...

lesswrong.com
lesswrong.com > posts > ixp9oJXzjA9LrwiZo > you-yes-you-need-a-february-2020-checklist-for-ai-policy

You (Yes, You) Need A February 2020 Checklist for AI Policy — LessWrong

12+ hour, 23+ min ago   (409+ words) TL;DR: You (Yes You) should prepare for a “February 2020” moment where suddenly AI policy becomes the most important issue in the world. You should be ready to take action if and when it does, in a detailed way. (Epistemic…...

lesswrong.com
lesswrong.com > posts > Rk6FbkDFFm8ciqefv > quadrillion-param-costs-kv-cache-context-length-frontier

Quadrillion Param Costs: KV Cache, Context Length, Frontier Margins — LessWrong

13+ hour, 5+ min ago   (448+ words) Before formulating how costs of tokens can be computed in general (which we'll need to consider their scaling, across different models and hardware systems), let's work through a concrete example of the 10T model of 2026 running on GB300 Oberon (NVL72). Each chip in…...