§ Reading · Field Radar
Field Radar.
What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.
How this list is made
This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.
The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.
Sources — LessWrong: error: HTTP 429 for https://www.lesswrong.com/graphql · Hacker News: ok (5 stories) · Reddit: ok (25 posts) — some subreddits rate-limited
- As of
- 2026-08-08 14:00 ET
- Showing
- 25 items
- New
- 0 in last 48h
- Refresh
- Every 6 hours
- 0.62LessWrong4dRLVR that rewards red teaming the training environment
Epistemic status: throwing an idea at the wall and seeing if it sticks I've been thinking about how to mitigate egregious reward hacking, a la the Hugging Face incident. I don't have the resources I'd need to write a…
why score 0.616
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.21 ×0.25 0.052 contributability 0.86 ×0.15 0.130 venue 0.85 ×0.10 0.085 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.55LessWrong2wA Multi-Agent Extension for Petri
Intro Petri is an open-source framework built on Inspect AI for automated AI Safety evaluations first released by Anthropic, but now maintained and developed by Meridian Labs. Each evaluation involves three agents, the…
why score 0.549
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.001 contributability 0.02 ×0.15 0.003 venue 0.45 ×0.10 0.045 direct 1.00 ×0.20 0.200 2 matching tag(s)
- 0.42Hacker News5dThe Download: reward hacking explained, and suspected Iranian cyberattacks
why score 0.420
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.19 ×0.25 0.047 contributability 0.02 ×0.15 0.003 venue 0.21 ×0.10 0.021 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.42LessWrong4dWhen should we trust a latent representation?
Working on AI Safety, I spend a significant part of my time thinking about how frontier AI Safety research eventually gets translated into deployable safety systems. One thing I have constantly noticed is the transition…
why score 0.419
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.24 ×0.25 0.060 contributability 0.12 ×0.15 0.017 venue 0.41 ×0.10 0.041 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.42LessWrong4dRewrite All the Code, All the Time
This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. It has become a mainstream prediction that…
why score 0.418
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.26 ×0.25 0.065 contributability 1.00 ×0.15 0.149 venue 0.53 ×0.10 0.053 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.42LessWrong2wFable is SOTA at CIFAR Speedrun (& specification gaming)
Fulcrum is working on an AI R&D optimization benchmark. Here, we present results from one of our tasks, including preliminary results from Fable. For more detail on Fable’s solution, check out…
why score 0.417
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.000 contributability 0.02 ×0.15 0.003 venue 0.64 ×0.10 0.064 direct 1.00 ×0.20 0.200 tier-1: specification gaming
- 0.41LessWrong5dSingle Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol
This is a research update for an on-going replication of single-forward-pass evals done as part of the Second Look Fellowship. In following posts, we will run more comprehensive replications of previous work and release…
why score 0.407
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.14 ×0.25 0.036 contributability 0.02 ×0.15 0.003 venue 0.69 ×0.10 0.069 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.40LessWrong4dDeliberate Alignment Faking as a Defense Against Model Poisoning
I want to discuss and brainstorm a counterintuitive approach to AI alignment: Inducing alignment faking on purpose, to prevent the model from developing emergent misalignment. To prevent this from going horribly wrong,…
why score 0.401
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.19 ×0.25 0.049 contributability 0.02 ×0.15 0.003 venue 0.50 ×0.10 0.050 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.40LessWrong11dMulti-Turn Drift Increases Scheming
TLDR - We talk about scheming, and why research on this phenomenon is crucial for AI safety. We find a particular environment/scenario where scheming happens at a higher rate than normal. We provide hypotheses for why…
why score 0.395
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.02 ×0.25 0.005 contributability 0.27 ×0.15 0.040 venue 0.51 ×0.10 0.051 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.38Hacker News2wShow HN: Vinv-Ties every runtime trace to code segment, prevents reward hacking
why score 0.380
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.01 ×0.25 0.002 contributability 0.02 ×0.15 0.003 venue 0.26 ×0.10 0.026 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.38Hacker News2wShow HN: VinvAI – Ties runtime trace to code segment, prevents reward hacking
why score 0.375
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.01 ×0.25 0.002 contributability 0.02 ×0.15 0.003 venue 0.21 ×0.10 0.021 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.37LessWrong2wLinear probes tell you where quantization will hurt
Epistemic status: I have only tested one encoder family (BERT-base and its relatives) and one decoder LLM (Qwen2.5-3B), one seed, token-level tasks, and post-training weight quantization. I trust the results because I…
why score 0.367
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.002 contributability 0.02 ×0.15 0.003 venue 0.63 ×0.10 0.063 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.36LessWrong4dConcrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face
Three-Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we describe the ambitious,…
why score 0.364
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.24 ×0.25 0.060 contributability 0.42 ×0.15 0.064 venue 0.90 ×0.10 0.090 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.36LessWrong2wV&V takes on OpenAI’s long-horizon incidents
[Cross-posted from The Foretellix CTO Blog. These short takes try to put a verification-and-validation slant on AI-safety / alignment topics – they are not full treatments. I co-originated coverage-driven verification…
why score 0.363
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.001 contributability 0.02 ×0.15 0.003 venue 0.59 ×0.10 0.059 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.36LessWrong2wIs there even a ground-truth for LLMs’ internal representations?
[This is an introductory blog for the paper Laguerre Geometry for Interpreting Large Language Models and the GitHub repository Geometric Lens.] LLM Lens: What does an internal vector mean? Anthropic's recent paper on…
why score 0.363
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.000 contributability 0.02 ×0.15 0.003 venue 0.60 ×0.10 0.060 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.36LessWrong2wThe State of AI Consciousness Research
Epistemic status: a survey, not an argument. I am agnostic on whether any current system is conscious; the claim is only that the question is researchable. This piece surveys the empirical research on AI consciousness.…
why score 0.362
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.00 ×0.25 0.000 contributability 0.42 ×0.15 0.064 venue 0.73 ×0.10 0.073 direct 0.00 ×0.20 0.000 tier-2: mechanistic interpretability; 1 matching tag(s)
- 0.36LessWrong2wSONI: Selective Orthogonalisation via Noise Injection
This project was completed as a capstone for TARA. All code is available in github. TL;DR The Problem: Neural networks use superposition to pack many concepts into small latent spaces by making feature vectors…
why score 0.362
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.002 contributability 0.02 ×0.15 0.003 venue 0.57 ×0.10 0.057 direct 0.00 ×0.20 0.000 tier-2: feature, activation; 1 matching tag(s)
- 0.36r/ControlProblem11dWhat failed in the Hugging Face incident wasn't the model
TL/DR: everyone is arguing about whether the Hugging Face models went rogue. They didn't, it's specification gaming, that argument is a decade old. The part being missed is that three separate checks sat above those…
why score 0.359
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.02 ×0.25 0.006 contributability 0.02 ×0.15 0.003 venue 0.00 ×0.10 0.000 direct 1.00 ×0.20 0.200 tier-1: specification gaming
- 0.36LessWrong2wMechanistic interpretability hypotheses for Measuring Reward-Seeking by Instilling Contrastive Beliefs and additional comments
This is interesting research! https://alignment.openai.com/measuring-reward-seeking It made me think of few overlapping hypotheses for what might be happening here, how did the grader behavior emerge at the pretraining…
why score 0.357
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.001 contributability 0.02 ×0.15 0.003 venue 0.53 ×0.10 0.053 direct 0.00 ×0.20 0.000 tier-2: mechanistic interpretability; 2 matching tag(s)
- 0.36r/ControlProblem13dIs specification downstream from judgment?
Hey everyone. I’ve long been fascinated by both philosophy of technology and AI alignment. I’m also using Heidegger quite a bit for my philosophy PhD. Given the recent OpenAI–Hugging Face incident reported this week, I…
why score 0.355
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.01 ×0.25 0.003 contributability 0.02 ×0.15 0.003 venue 0.00 ×0.10 0.000 direct 1.00 ×0.20 0.200 tier-1: specification gaming
- 0.36LessWrong12dLLMs are (still) mostly powered by imitative learning, not RL
Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It’s so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture. Stepping back,…
why score 0.355
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.01 ×0.25 0.003 contributability 0.72 ×0.15 0.108 venue 0.94 ×0.10 0.094 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.35LessWrong12dInoculate or Reflect? Two training interventions under prompting, steering, and patching
Anthropic's recent paper, Verbalizable Representations Form a Global Workspace in Language Models, contains a small experiment near the end that we found more interesting than the main findings. Surprising that it's so…
why score 0.349
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.003 contributability 0.02 ×0.15 0.003 venue 0.43 ×0.10 0.043 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.35LessWrong2wSteering Blackmail Through a Model's "Emotional State"
Gemma 3 12B's choice to blackmail is visible inside the model, but the obvious signal is not a useful control switch; surprisingly, a nearby "desperate versus calm" signal is. In this case study, I show how that…
why score 0.348
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.001 contributability 0.02 ×0.15 0.003 venue 0.45 ×0.10 0.045 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.35LessWrong2wTracing causal structure in LLM-generated text: a different lens on the Dallas circuit
The classic "Dallas" example from Anthropic focuses on an internal circuit in an LLM. I became curious about what the same underlying process looks like when viewed through the generated reasoning trace instead of…
why score 0.348
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.000 contributability 0.02 ×0.15 0.003 venue 0.45 ×0.10 0.045 direct 0.00 ×0.20 0.000 tier-2: circuit; 2 matching tag(s)
- 0.35LessWrong11dWhen the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
By Sai Kartheek Reddy Kasu, Nils Lukas, and Samuele Poppi This post is a summary of our accepted paper at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN). The full paper is available here TL;DR The Setup:…
why score 0.346
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.02 ×0.25 0.005 contributability 0.02 ×0.15 0.003 venue 0.39 ×0.10 0.039 direct 0.00 ×0.20 0.000 2 matching tag(s)