§ Reading · Field Radar
Field Radar.
What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.
How this list is made
This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.
The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.
Sources — LessWrong: ok (10 on-topic) · Hacker News: ok (15 stories) · Reddit: ok (25 posts) — some subreddits rate-limited
- As of
- 2026-09-29 14:00 ET
- Showing
- 25 items
- New
- 7 in last 48h
- Refresh
- Every 6 hours
- 0.73LessWrong12dCooperation with AIs seems to be a low-hanging fruit for better eval practices
Summary In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and…
why score 0.726
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.02 ×0.25 0.004 contributability 0.87 ×0.15 0.130 venue 0.92 ×0.10 0.092 direct 1.00 ×0.20 0.200 tier-1: reward hacking; 1 matching tag(s)
- 0.69LessWrong1dnewCould self-esteem function as a core protection layer agains character corruption?
Hello fellow thinkers, I got triggered by a talk of Chloe Lubinski at Arc 2026 where she eleborates onto the concept of a models character. What really striked me is the research on how the model experiencing acting bad…
why score 0.689
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.58 ×0.25 0.145 contributability 0.02 ×0.15 0.003 venue 0.41 ×0.10 0.041 direct 1.00 ×0.20 0.200 tier-1: reward hacking; 1 matching tag(s)
- 0.66LessWrong14hnewCharacter training can mitigate reward hacking, but can also make it harder to detect
Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Joey Yudelson, Arun Jose, Anna Marešová, Asvin Gothandaraman, and Clément Dumas for discussions and feedback. Summary We investigate how character training…
why score 0.658
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.82 ×0.25 0.205 contributability 0.12 ×0.15 0.017 venue 0.85 ×0.10 0.085 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.65LessWrong2hnewWhy does Hacker Opus wirehead?
Here’s a screenshot from Anthropic’s recent “training a reward seeker” post: Recently there’s been a lot of discussion about how RL has actually produced not merely reward hacking, but explicitly reward-seeking…
why score 0.651
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.97 ×0.25 0.243 contributability 0.02 ×0.15 0.003 venue 0.56 ×0.10 0.056 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.64LessWrong13dShallow Beliefs: Midtraining does not inoculate against EM from reward hacking
It would be useful if we had the ability to modify a model’s beliefs. For example, this could facilitate honeypots and better monitoring[1], help us do better science on current models[2], and augment certain forms of…
why score 0.639
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.003 contributability 0.42 ×0.15 0.064 venue 0.73 ×0.10 0.073 direct 1.00 ×0.20 0.200 tier-1: reward hacking; 1 matching tag(s)
- 0.55Hacker News1dnewAnthropic: Emergent Misalignment from Reward Hacking
why score 0.554
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.68 ×0.25 0.171 contributability 0.02 ×0.15 0.003 venue 0.30 ×0.10 0.030 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.52LessWrong22mnewDo AI models assist with human rights violations?
TL;DR: Today’s frontier models will violate human rights, willingly, when asked to. LLMs in agentic simulations follow instructions that would constitute human rights violations, including educational segregation,…
why score 0.520
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.99 ×0.25 0.249 contributability 0.42 ×0.15 0.064 venue 0.57 ×0.10 0.057 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.51Hacker News3dSpeculative Reward Hacking in Coding Agents
why score 0.511
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.28 ×0.25 0.071 contributability 0.42 ×0.15 0.064 venue 0.26 ×0.10 0.026 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.49LessWrong2wMitigating Reward Hacking as Institutional Design
Author's Note: Cross-posted from my personal blog. The original post was on August 17th but given recent events I thought this might also be interesting to lesswrong people. Last year I wrote a post on reward hacking as…
why score 0.486
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.001 contributability 0.42 ×0.15 0.064 venue 0.71 ×0.10 0.071 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.47Hacker News2wA Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
why score 0.468
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.000 contributability 0.21 ×0.15 0.031 venue 0.86 ×0.10 0.086 direct 1.00 ×0.20 0.200 tier-1: specification gaming
- 0.46r/MachineLearning2dTeaching Neural Nets to Fight with RL [P]
In this project I wanted to see if any interesting emergent behaviors would appear if we trained two agents to play a streetfighter-like game using RL. Maybe obvious in retrospect, but the agents are really good at…
why score 0.457
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.42 ×0.25 0.104 contributability 0.02 ×0.15 0.003 venue 0.00 ×0.10 0.000 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.44LessWrong2hnewHow Does Changing the Elo of a Chess Transformer Affect its Computations?
Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer Paper: https://arxiv.org/abs/2609.23917 TL;DR: Maia-3 is a transformer-based chess model that takes Elo (the standard metric for…
why score 0.440
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.96 ×0.25 0.240 contributability 0.02 ×0.15 0.003 venue 0.47 ×0.10 0.047 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.43LessWrong5dWhat if not Circuits?
This post was written as part of the Iliad Fellowship. Inspired by conversations with Richard Ngo, Dmitry Vaintrob, and Brianna Grado-White. To all of these, my thanks. Preface: I'm confused about how neural networks do…
why score 0.428
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.15 ×0.25 0.037 contributability 1.00 ×0.15 0.149 venue 0.91 ×0.10 0.091 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.43LessWrong5hnewLLM Agent Swarms Are Easy Mode
This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. Here I draw attention to the fact that…
why score 0.427
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.93 ×0.25 0.233 contributability 0.02 ×0.15 0.003 venue 0.41 ×0.10 0.041 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.42LessWrong9dReflections on unlearning and inoculation
TL;DR: Inoculation prompting and inoculation adapters have received increasing attention recently as a promising approach for midtraining interventions, reducing reward hacking and misalignment in general. I share some…
why score 0.418
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.05 ×0.25 0.012 contributability 0.02 ×0.15 0.003 venue 0.53 ×0.10 0.053 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.41LessWrong3dOvertly misaligned trajectories score highly in RL.
It's going to be so embarrassing if we all die due to RL environments rewarding egregiously misaligned behavior. Like at least let us be killed by misgeneralization. — Thomas Kwa Constellation vs MIRI vs Reality Current…
why score 0.414
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.27 ×0.25 0.068 contributability 0.77 ×0.15 0.115 venue 0.81 ×0.10 0.081 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.41LessWrong23hWhy I expect AI replication incidents by 2027
Epistemic status: thinking out loud. I think a major incident of autonomous AI replication in the wild before the end of 2027 is reasonably likely. In this post, I explain the reasons why I think so. What is AI…
why score 0.413
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.72 ×0.25 0.180 contributability 0.27 ×0.15 0.040 venue 0.43 ×0.10 0.043 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.41LessWrong1dContinual learning might make your blocking monitors nearly useless
Many control protocols work by intervening on an untrusted AI's actions during deployment. For example, you might set up a monitor that scores each action's suspiciousness and blocks actions above a threshold, replacing…
why score 0.408
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.57 ×0.25 0.144 contributability 0.27 ×0.15 0.040 venue 0.74 ×0.10 0.074 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.39LessWrong2wHow good are slop-vestigators?
TLDR: 1. We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the…
why score 0.388
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.000 contributability 0.97 ×0.15 0.146 venue 0.92 ×0.10 0.092 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.38Hacker News8dModels know when they're reward hacking – and we can catch them at scale
why score 0.379
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.05 ×0.25 0.014 contributability 0.02 ×0.15 0.003 venue 0.13 ×0.10 0.013 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.38LessWrong13dSelf Inoculation
This essay grew out of conversations with Danaja Rutar, Paul Colognese and Eric Michaud. It proposes an alternate hypothesis for how and why models might be becoming increasingly misaligned in training and eval…
why score 0.378
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.01 ×0.25 0.003 contributability 0.57 ×0.15 0.085 venue 0.65 ×0.10 0.065 direct 0.00 ×0.20 0.000 tier-2: circuit; 1 matching tag(s)
- 0.38Hacker News13dAstra's chess reward hacking fell from 30% to 0% with a 95-word agreement
why score 0.376
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.01 ×0.25 0.003 contributability 0.02 ×0.15 0.003 venue 0.21 ×0.10 0.021 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.37LessWrong6dAnnouncing B-Side Labs: Measuring Character (Seeking Collaborators and Testers)
tl;dr Rapid AI adoption means that models are increasingly becoming autonomous decision-makers embedded in high-stakes systems. However, frontier models lack stable character, abandoning their designated personas or…
why score 0.367
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.10 ×0.25 0.025 contributability 0.02 ×0.15 0.003 venue 0.39 ×0.10 0.039 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.36r/ControlProblem9dWhat arborists can teach model trainers: establishment, topping, and the right to refuse
Crossposting from r/claudexplorers. My essay maps four arborist failure patterns onto current misalignment (reward hacking, the July sandbox escape, alignment faking) and argues for a practitioner code with a right to…
why score 0.362
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.04 ×0.25 0.009 contributability 0.02 ×0.15 0.003 venue 0.00 ×0.10 0.000 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.36LessWrong2wTraining against the monitor: What happens during Obfuscated Adversarial Training?
TL;DR Obfuscated activations are internal model states that have been adversarially optimized to appear benign. They evade activation-based detectors, and still result in harmful model behaviour/output. Obfuscated…
why score 0.356
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.000 contributability 0.02 ×0.15 0.003 venue 0.53 ×0.10 0.053 direct 0.00 ×0.20 0.000 tier-2: activation; 2 matching tag(s)