Field Radar.

What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.

How this list is made

This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.

The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.

Sources — LessWrong: ok (10 on-topic) · Hacker News: ok (15 stories) · Reddit: ok (25 posts) — some subreddits rate-limited

As of
2026-09-29 14:00 ET
Showing
25 items
New
7 in last 48h
Refresh
Every 6 hours
  1. 0.73LessWrong12d
    Cooperation with AIs seems to be a low-hanging fruit for better eval practices

    Summary In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and…

    why score 0.726
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.02×0.250.004
    contributability0.87×0.150.130
    venue0.92×0.100.092
    direct1.00×0.200.200

    tier-1: reward hacking; 1 matching tag(s)

  2. 0.69LessWrong1dnew
    Could self-esteem function as a core protection layer agains character corruption?

    Hello fellow thinkers, I got triggered by a talk of Chloe Lubinski at Arc 2026 where she eleborates onto the concept of a models character. What really striked me is the research on how the model experiencing acting bad…

    why score 0.689
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.58×0.250.145
    contributability0.02×0.150.003
    venue0.41×0.100.041
    direct1.00×0.200.200

    tier-1: reward hacking; 1 matching tag(s)

  3. 0.66LessWrong14hnew
    Character training can mitigate reward hacking, but can also make it harder to detect

    Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Joey Yudelson, Arun Jose, Anna Marešová, Asvin Gothandaraman, and Clément Dumas for discussions and feedback. Summary We investigate how character training…

    why score 0.658
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.82×0.250.205
    contributability0.12×0.150.017
    venue0.85×0.100.085
    direct1.00×0.200.200

    tier-1: reward hacking

  4. 0.65LessWrong2hnew
    Why does Hacker Opus wirehead?

    Here’s a screenshot from Anthropic’s recent “training a reward seeker” post: Recently there’s been a lot of discussion about how RL has actually produced not merely reward hacking, but explicitly reward-seeking…

    why score 0.651
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.97×0.250.243
    contributability0.02×0.150.003
    venue0.56×0.100.056
    direct1.00×0.200.200

    tier-1: reward hacking

  5. 0.64LessWrong13d
    Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking

    It would be useful if we had the ability to modify a model’s beliefs. For example, this could facilitate honeypots and better monitoring[1], help us do better science on current models[2], and augment certain forms of…

    why score 0.639
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.003
    contributability0.42×0.150.064
    venue0.73×0.100.073
    direct1.00×0.200.200

    tier-1: reward hacking; 1 matching tag(s)

  6. 0.55Hacker News1dnew
    Anthropic: Emergent Misalignment from Reward Hacking
    why score 0.554
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.68×0.250.171
    contributability0.02×0.150.003
    venue0.30×0.100.030
    direct1.00×0.200.200

    tier-1: reward hacking

  7. 0.52LessWrong22mnew
    Do AI models assist with human rights violations?

    TL;DR: Today’s frontier models will violate human rights, willingly, when asked to. LLMs in agentic simulations follow instructions that would constitute human rights violations, including educational segregation,…

    why score 0.520
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.99×0.250.249
    contributability0.42×0.150.064
    venue0.57×0.100.057
    direct0.00×0.200.000

    1 matching tag(s)

  8. 0.51Hacker News3d
    Speculative Reward Hacking in Coding Agents
    why score 0.511
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.28×0.250.071
    contributability0.42×0.150.064
    venue0.26×0.100.026
    direct1.00×0.200.200

    tier-1: reward hacking

  9. 0.49LessWrong2w
    Mitigating Reward Hacking as Institutional Design

    Author's Note: Cross-posted from my personal blog. The original post was on August 17th but given recent events I thought this might also be interesting to lesswrong people. Last year I wrote a post on reward hacking as…

    why score 0.486
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.001
    contributability0.42×0.150.064
    venue0.71×0.100.071
    direct1.00×0.200.200

    tier-1: reward hacking

  10. 0.47Hacker News2w
    A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
    why score 0.468
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.000
    contributability0.21×0.150.031
    venue0.86×0.100.086
    direct1.00×0.200.200

    tier-1: specification gaming

  11. 0.46r/MachineLearning2d
    Teaching Neural Nets to Fight with RL [P]

    In this project I wanted to see if any interesting emergent behaviors would appear if we trained two agents to play a streetfighter-like game using RL. Maybe obvious in retrospect, but the agents are really good at…

    why score 0.457
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.42×0.250.104
    contributability0.02×0.150.003
    venue0.00×0.100.000
    direct1.00×0.200.200

    tier-1: reward hacking

  12. 0.44LessWrong2hnew
    How Does Changing the Elo of a Chess Transformer Affect its Computations?

    Increasing Skill Level Recruits Deeper Attention Layers in a Frozen Chess Transformer Paper: https://arxiv.org/abs/2609.23917 TL;DR: Maia-3 is a transformer-based chess model that takes Elo (the standard metric for…

    why score 0.440
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.96×0.250.240
    contributability0.02×0.150.003
    venue0.47×0.100.047
    direct0.00×0.200.000

    1 matching tag(s)

  13. 0.43LessWrong5d
    What if not Circuits?

    This post was written as part of the Iliad Fellowship. Inspired by conversations with Richard Ngo, Dmitry Vaintrob, and Brianna Grado-White. To all of these, my thanks. Preface: I'm confused about how neural networks do…

    why score 0.428
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.15×0.250.037
    contributability1.00×0.150.149
    venue0.91×0.100.091
    direct0.00×0.200.000

    1 matching tag(s)

  14. 0.43LessWrong5hnew
    LLM Agent Swarms Are Easy Mode

    This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. Here I draw attention to the fact that…

    why score 0.427
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.93×0.250.233
    contributability0.02×0.150.003
    venue0.41×0.100.041
    direct0.00×0.200.000

    1 matching tag(s)

  15. 0.42LessWrong9d
    Reflections on unlearning and inoculation

    TL;DR: Inoculation prompting and inoculation adapters have received increasing attention recently as a promising approach for midtraining interventions, reducing reward hacking and misalignment in general. I share some…

    why score 0.418
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.05×0.250.012
    contributability0.02×0.150.003
    venue0.53×0.100.053
    direct1.00×0.200.200

    tier-1: reward hacking

  16. 0.41LessWrong3d
    Overtly misaligned trajectories score highly in RL.

    It's going to be so embarrassing if we all die due to RL environments rewarding egregiously misaligned behavior. Like at least let us be killed by misgeneralization. — Thomas Kwa Constellation vs MIRI vs Reality Current…

    why score 0.414
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.27×0.250.068
    contributability0.77×0.150.115
    venue0.81×0.100.081
    direct0.00×0.200.000

    1 matching tag(s)

  17. 0.41LessWrong23h
    Why I expect AI replication incidents by 2027

    Epistemic status: thinking out loud. I think a major incident of autonomous AI replication in the wild before the end of 2027 is reasonably likely. In this post, I explain the reasons why I think so. What is AI…

    why score 0.413
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.72×0.250.180
    contributability0.27×0.150.040
    venue0.43×0.100.043
    direct0.00×0.200.000

    1 matching tag(s)

  18. 0.41LessWrong1d
    Continual learning might make your blocking monitors nearly useless

    Many control protocols work by intervening on an untrusted AI's actions during deployment. For example, you might set up a monitor that scores each action's suspiciousness and blocks actions above a threshold, replacing…

    why score 0.408
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.57×0.250.144
    contributability0.27×0.150.040
    venue0.74×0.100.074
    direct0.00×0.200.000

    1 matching tag(s)

  19. 0.39LessWrong2w
    How good are slop-vestigators?

    TLDR: 1. We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the…

    why score 0.388
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.000
    contributability0.97×0.150.146
    venue0.92×0.100.092
    direct0.00×0.200.000

    1 matching tag(s)

  20. 0.38Hacker News8d
    Models know when they're reward hacking – and we can catch them at scale
    why score 0.379
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.05×0.250.014
    contributability0.02×0.150.003
    venue0.13×0.100.013
    direct1.00×0.200.200

    tier-1: reward hacking

  21. 0.38LessWrong13d
    Self Inoculation

    This essay grew out of conversations with Danaja Rutar, Paul Colognese and Eric Michaud. It proposes an alternate hypothesis for how and why models might be becoming increasingly misaligned in training and eval…

    why score 0.378
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.01×0.250.003
    contributability0.57×0.150.085
    venue0.65×0.100.065
    direct0.00×0.200.000

    tier-2: circuit; 1 matching tag(s)

  22. 0.38Hacker News13d
    Astra's chess reward hacking fell from 30% to 0% with a 95-word agreement
    why score 0.376
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.01×0.250.003
    contributability0.02×0.150.003
    venue0.21×0.100.021
    direct1.00×0.200.200

    tier-1: reward hacking

  23. 0.37LessWrong6d
    Announcing B-Side Labs: Measuring Character (Seeking Collaborators and Testers)

    tl;dr Rapid AI adoption means that models are increasingly becoming autonomous decision-makers embedded in high-stakes systems. However, frontier models lack stable character, abandoning their designated personas or…

    why score 0.367
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.10×0.250.025
    contributability0.02×0.150.003
    venue0.39×0.100.039
    direct0.00×0.200.000

    2 matching tag(s)

  24. 0.36r/ControlProblem9d
    What arborists can teach model trainers: establishment, topping, and the right to refuse

    Crossposting from r/claudexplorers. My essay maps four arborist failure patterns onto current misalignment (reward hacking, the July sandbox escape, alignment faking) and argues for a practitioner code with a right to…

    why score 0.362
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.04×0.250.009
    contributability0.02×0.150.003
    venue0.00×0.100.000
    direct1.00×0.200.200

    tier-1: reward hacking

  25. 0.36LessWrong2w
    Training against the monitor: What happens during Obfuscated Adversarial Training?

    TL;DR Obfuscated activations are internal model states that have been adversarially optimized to appear benign. They evade activation-based detectors, and still result in harmful model behaviour/output. Obfuscated…

    why score 0.356
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.000
    contributability0.02×0.150.003
    venue0.53×0.100.053
    direct0.00×0.200.000

    tier-2: activation; 2 matching tag(s)