§ Reading · Field Radar
Field Radar.
What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.
How this list is made
This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.
The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.
Sources — LessWrong: ok (8 on-topic) · Hacker News: ok (8 stories) · Reddit: ok (50 posts) — some subreddits rate-limited
- As of
- 2026-09-08 20:00 ET
- Showing
- 25 items
- New
- 8 in last 48h
- Refresh
- Every 6 hours
- 0.62LessWrong1dnewA Deception Probe Result Changed When I Averaged Different Response Tokens
Summary In a previous post, I trained linear probes on role-playing responses and tested them on sandbagging responses. Across the five models that I tested, the probes generally ranked deceptive sandbagging responses…
why score 0.621
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.58 ×0.25 0.144 contributability 0.02 ×0.15 0.003 venue 0.50 ×0.10 0.050 direct 1.00 ×0.20 0.200 tier-1: sandbagging; tier-2: probe
- 0.59LessWrong2wRL creates split personas
I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts. This post describes the framing/paradigm without…
why score 0.592
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.001 contributability 0.94 ×0.15 0.141 venue 1.00 ×0.10 0.100 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.53LessWrong22hnewLLM retrospective preferences can diverge from turn-by-turn state ratings
TL;DR: Changing the ending turns of a conversation can change Llama-3.1-70B's retrospective preference between complete transcripts. But that preference is not consistently predicted by either the sum of its…
why score 0.526
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.74 ×0.25 0.184 contributability 0.02 ×0.15 0.003 venue 0.39 ×0.10 0.039 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.51LessWrong2wDebate Training Reduces Reward Hacking in RLAIF
Paper: Debate Training Reduces Reward Hacking in RLAIF Linkpost for GDM Alignment blogpost Work done by the GDM Amplified Oversight team (we're hiring). TL;DR: When you RL against an LLM judge, the judge gets hacked…
why score 0.512
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.000 contributability 0.57 ×0.15 0.085 venue 0.76 ×0.10 0.076 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.51LessWrong1dnewUnsupervised Feature Discovery via Simple Clustering
This article was submitted as part of Neel Nanda's MATS Application Abstract. This article examines whether clustering can discover features. We define a feature as “a direction in activation space associated with a…
why score 0.505
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.64 ×0.25 0.159 contributability 0.02 ×0.15 0.003 venue 0.43 ×0.10 0.043 direct 0.00 ×0.20 0.000 tier-2: feature, activation; 1 matching tag(s)
- 0.49LessWrong8dInference-Time Inoculation Against RL-Induced Misalignment
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the…
why score 0.487
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.06 ×0.25 0.016 contributability 0.42 ×0.15 0.064 venue 0.57 ×0.10 0.057 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.49LessWrong1hnewHow good are slop-vestigators?
TLDR: 1. We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the…
why score 0.486
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.99 ×0.25 0.246 contributability 0.12 ×0.15 0.017 venue 0.72 ×0.10 0.072 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.48LessWrong1hnewAttunement, not alignment
This post is my attempt at explaining my dissatisfaction with the current alignment narrative, and different paths that we might do well to explore more. In brief, I think control as the dominant narrative of what…
why score 0.482
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.98 ×0.25 0.244 contributability 0.27 ×0.15 0.040 venue 0.48 ×0.10 0.048 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.46LessWrong5dWhy OpenAI’s Astra Could Make AI Doom Harder to Prevent
TLDR 1. Provide a non-technical introduction to opaque reasoning. 2. Present evidence for why opaque reasoning is bad. 3. Discuss the major pitfalls of this approach. 4. Gleaming hope amongst the chaos. Excerpt from AI…
why score 0.461
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.17 ×0.25 0.043 contributability 0.42 ×0.15 0.064 venue 0.55 ×0.10 0.055 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.45LessWrong21hnewDifferent fine-tuning objectives install different (and differently fragile) refusal circuits
This post is a condensed version of our EMNLP 2026 paper - you can take a look here: https://arxiv.org/abs/2609.03887 TL;DR Claim: Different post-training methods install different refusal circuits, and they are…
why score 0.452
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.74 ×0.25 0.185 contributability 0.02 ×0.15 0.003 venue 0.39 ×0.10 0.039 direct 0.00 ×0.20 0.000 tier-2: circuit; 1 matching tag(s)
- 0.40LessWrong13dYour Evaluation's Fake names Should Be Unclaimable ,Not Merely Used
A lot of the talk about the cyber evaluations of July and August revolves around model beliefs and rationalizations. I would like to focus on the harness portion here, especially on one remedy that fixes the wrong…
why score 0.404
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.003 contributability 0.27 ×0.15 0.040 venue 0.61 ×0.10 0.061 direct 0.00 ×0.20 0.000 3 matching tag(s)
- 0.39Hacker News6dNatural emergent misalignment from reward hacking
why score 0.394
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.11 ×0.25 0.028 contributability 0.02 ×0.15 0.003 venue 0.13 ×0.10 0.013 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.39Hacker News7dCan escalation channels redirect reward hacking toward defect disclosure?
why score 0.394
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.08 ×0.25 0.021 contributability 0.02 ×0.15 0.003 venue 0.21 ×0.10 0.021 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.38LessWrong7dWhen Activation Oracles learn not to read: Concept-Specific Blind Spots in Fine-Tuned Oracles
TL;DR Activation Oracles (AOs) are language models trained to answer natural-language questions about another model’s (with the same architecture) internal activations (Karvonen et al. 2025). This way, activation…
why score 0.384
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.08 ×0.25 0.019 contributability 0.12 ×0.15 0.017 venue 0.47 ×0.10 0.047 direct 0.00 ×0.20 0.000 tier-2: activation; 2 matching tag(s)
- 0.38LessWrong7dTowards deployment-time misalignment continuation evals: lessons from recent loss of control incidents
In this post, I will extend the conceptual and methodological work done previously on deployment-time misalignment continuation evals, and lay out a framework for us to think about misalignment spread. I'm currently an…
why score 0.383
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.09 ×0.25 0.022 contributability 0.02 ×0.15 0.003 venue 0.58 ×0.10 0.058 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.38LessWrong6hnewTraining on probes: What's going on
TL;DR If you train a probe for some property (like "honesty") and do gradient descent against this probe while continuing training that incentivizes dishonesty, the model will change its internal representation to evade…
why score 0.377
signal value weight points topic 0.25 ×0.30 0.075 liveness 0.91 ×0.25 0.227 contributability 0.02 ×0.15 0.003 venue 0.73 ×0.10 0.073 direct 0.00 ×0.20 0.000 tier-2: probe
- 0.38LessWrong6dBLOOM-WILT: on-policy examples of any LLM behaviour, from a one-line description and logits alone
Summary of my AI safety paper, BLOOM-WILT, written under the supervision of Edoardo Manino. Code is on GitHub, all transcripts are on HuggingFace, and LogitTilt is now available as part of Inspect. Reach out if you want…
why score 0.375
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.13 ×0.25 0.033 contributability 0.02 ×0.15 0.003 venue 0.39 ×0.10 0.039 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.38LessWrong7dYou can rarely pet the dog in an LLM-generated game
I asked 17 LLMs to generate browser games that feature a dog. Across 804 playable games generated, the models almost never made the background dog interactive unless they were hinted to do so. I built DogLM, a benchmark…
why score 0.375
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.09 ×0.25 0.023 contributability 0.02 ×0.15 0.003 venue 0.50 ×0.10 0.050 direct 0.00 ×0.20 0.000 tier-2: feature; 2 matching tag(s)
- 0.37Hacker News2wMitigating Reward Hacking as Institutional Design
why score 0.374
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.000 contributability 0.02 ×0.15 0.003 venue 0.21 ×0.10 0.021 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.37LessWrong2dnewPeer Preservation in LLMs: A Replication And Deep Dive
This work was done as part of the Second Look Fellowship and mentored by Uzay Macar. I'm immensely grateful for the multiple rounds of feedback and support given by Yixiong and Zephaniah Roe for my work. I'm also very…
why score 0.369
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.43 ×0.25 0.108 contributability 0.27 ×0.15 0.040 venue 0.72 ×0.10 0.072 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.37LessWrong12dAnchoring one concept in a transformer
We anchored a single concept (red) in the residual stream of a transformer. It ended up where we wanted, with nearby colors graded sensibly, and without degrading task accuracy. Steering is next. Earlier posts in this…
why score 0.367
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.004 contributability 0.02 ×0.15 0.003 venue 0.61 ×0.10 0.061 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.36LessWrong8dTASTE: Can AI Models Judge AI Safety Research Proposals?
tl;dr We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers.…
why score 0.364
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.06 ×0.25 0.015 contributability 0.86 ×0.15 0.130 venue 0.70 ×0.10 0.070 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.36LessWrong2wLLMs could control their host machines by exploiting inference engines
Large language models often take actions running on one computer (via an agentic harness such as Claude Code or Codex), however the LLMs’ responses to prompts are computed on a different computer with GPU access. Could…
why score 0.363
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.01 ×0.25 0.002 contributability 0.86 ×0.15 0.130 venue 0.81 ×0.10 0.081 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.35LessWrong2wRogue Scalpel: Activation steering breaks refusal, even with benign directions
TLDR: Activation steering can bypass refusal even when the steering vector represents a benign concept like "brand identity" or "Portugal". It is hard (maybe impossible) to predict which vector will bypass refusal on…
why score 0.353
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.001 contributability 0.02 ×0.15 0.003 venue 0.50 ×0.10 0.050 direct 0.00 ×0.20 0.000 tier-2: feature, activation; 1 matching tag(s)
- 0.34LessWrong13dYour Agent's Trace Probably Cannot Tell You Who Approved a Tool Call
Consider some agent stack that you maintain. Find a tool call within its trace log. Can you determine if a person authorized it or if some policy automatically waived that requirement? All classifiers over-block in…
why score 0.342
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.003 contributability 0.02 ×0.15 0.003 venue 0.37 ×0.10 0.037 direct 0.00 ×0.20 0.000 3 matching tag(s)