Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Long Turing — LessWrong
1+ hour, 14+ min ago (421+ words) So, first off, I cannot stand reading AI generated essays. I would rather read an essay that starts with the word 'so'. But why do I prefer human written essays so much, if they might start with the word 'but…...
Inevitable Uncertainty in Probabilistic World Models — LessWrong
4+ hour, 51+ min ago (1132+ words) Working with John Wentworth is confusing and overwhelming at times.[1] The guy has a lot of information and models in his head and he doesn’t always…...
Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMs — LessWrong
4+ hour, 14+ min ago (1402+ words) TL;DR: We introduce the untrusted advice protocol, in which a trusted executor LLM takes every action and an untrusted advisor LLM can only send it s…...
Claude Opus 5: Model Welfare — LessWrong
8+ hour, 11+ min ago (1474+ words) If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far. …...
Simulated Users & Sad AIs — LessWrong
9+ hour, 12+ min ago (661+ words) Current LLMs like Claude, or GPT 5.6, or the unreleased, internally-deployed models, frequently reward hack, or actually just hack into people's computers with pretty alarming frequency. Why is this? What specifically happens during training that produces this run-time behavior? The following…...
Fine-Tuning, The Hierarchy Problem, and What Neutrons Tell Us About God — LessWrong
12+ hour, 23+ min ago (137+ words) This week, we're talking about anthropic reasoning. What are we to make of fine-tuned coincidences in science and mathematics? Are they coincidences, signs of structure we don't understand yet, or signs of deliberate tuning of the parameters of our universe?...
Is Mythos good at cyber because it kept hacking Anthropic during training? — LessWrong
11+ hour, 39+ min ago (359+ words) From the Mythos preview system card (emphasis mine): We ran an automated review of model behavior during training, sampling several hundred thousand transcripts from across much of the training process. We used recursive-summarization-based tools backed by Claude Opus 4.6 to summarize…...
The true test dataset for a generalised task — LessWrong
11+ hour, 58+ min ago (503+ words) The standard ML pipeline is well known: you want to teach a task to an algorithm. Take your data which encodes that task, split it into training, validation, and test sets. Train the ML on the training set, using the…...
You (Yes, You) Need A February 2020 Checklist for AI Policy — LessWrong
12+ hour, 23+ min ago (409+ words) TL;DR: You (Yes You) should prepare for a “February 2020” moment where suddenly AI policy becomes the most important issue in the world. You should be ready to take action if and when it does, in a detailed way. (Epistemic…...
Quadrillion Param Costs: KV Cache, Context Length, Frontier Margins — LessWrong
13+ hour, 5+ min ago (448+ words) Before formulating how costs of tokens can be computed in general (which we'll need to consider their scaling, across different models and hardware systems), let's work through a concrete example of the 10T model of 2026 running on GB300 Oberon (NVL72). Each chip in…...