Field Radar.

What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.

How this list is made

This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.

The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.

Sources — LessWrong: ok (8 on-topic) · Hacker News: ok (8 stories) · Reddit: ok (50 posts) — some subreddits rate-limited

As of
2026-09-08 20:00 ET
Showing
25 items
New
8 in last 48h
Refresh
Every 6 hours
  1. 0.62LessWrong1dnew
    A Deception Probe Result Changed When I Averaged Different Response Tokens

    Summary In a previous post, I trained linear probes on role-playing responses and tested them on sandbagging responses. Across the five models that I tested, the probes generally ranked deceptive sandbagging responses…

    why score 0.621
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.58×0.250.144
    contributability0.02×0.150.003
    venue0.50×0.100.050
    direct1.00×0.200.200

    tier-1: sandbagging; tier-2: probe

  2. 0.59LessWrong2w
    RL creates split personas

    I describe my current view of personas in LLMs and why RL leads to egregious reward hacking in some contexts while the same models seem very aligned in other contexts. This post describes the framing/paradigm without…

    why score 0.592
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.001
    contributability0.94×0.150.141
    venue1.00×0.100.100
    direct1.00×0.200.200

    tier-1: reward hacking

  3. 0.53LessWrong22hnew
    LLM retrospective preferences can diverge from turn-by-turn state ratings

    TL;DR: Changing the ending turns of a conversation can change Llama-3.1-70B's retrospective preference between complete transcripts. But that preference is not consistently predicted by either the sum of its…

    why score 0.526
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.74×0.250.184
    contributability0.02×0.150.003
    venue0.39×0.100.039
    direct0.00×0.200.000

    2 matching tag(s)

  4. 0.51LessWrong2w
    Debate Training Reduces Reward Hacking in RLAIF

    Paper: Debate Training Reduces Reward Hacking in RLAIF Linkpost for GDM Alignment blogpost Work done by the GDM Amplified Oversight team (we're hiring). TL;DR: When you RL against an LLM judge, the judge gets hacked…

    why score 0.512
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.000
    contributability0.57×0.150.085
    venue0.76×0.100.076
    direct1.00×0.200.200

    tier-1: reward hacking

  5. 0.51LessWrong1dnew
    Unsupervised Feature Discovery via Simple Clustering

    This article was submitted as part of Neel Nanda's MATS Application Abstract. This article examines whether clustering can discover features. We define a feature as “a direction in activation space associated with a…

    why score 0.505
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.64×0.250.159
    contributability0.02×0.150.003
    venue0.43×0.100.043
    direct0.00×0.200.000

    tier-2: feature, activation; 1 matching tag(s)

  6. 0.49LessWrong8d
    Inference-Time Inoculation Against RL-Induced Misalignment

    Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the…

    why score 0.487
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.06×0.250.016
    contributability0.42×0.150.064
    venue0.57×0.100.057
    direct1.00×0.200.200

    tier-1: reward hacking

  7. 0.49LessWrong1hnew
    How good are slop-vestigators?

    TLDR: 1. We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the…

    why score 0.486
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.99×0.250.246
    contributability0.12×0.150.017
    venue0.72×0.100.072
    direct0.00×0.200.000

    1 matching tag(s)

  8. 0.48LessWrong1hnew
    Attunement, not alignment

    This post is my attempt at explaining my dissatisfaction with the current alignment narrative, and different paths that we might do well to explore more. In brief, I think control as the dominant narrative of what…

    why score 0.482
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.98×0.250.244
    contributability0.27×0.150.040
    venue0.48×0.100.048
    direct0.00×0.200.000

    1 matching tag(s)

  9. 0.46LessWrong5d
    Why OpenAI’s Astra Could Make AI Doom Harder to Prevent

    TLDR 1. Provide a non-technical introduction to opaque reasoning. 2. Present evidence for why opaque reasoning is bad. 3. Discuss the major pitfalls of this approach. 4. Gleaming hope amongst the chaos. Excerpt from AI…

    why score 0.461
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.17×0.250.043
    contributability0.42×0.150.064
    venue0.55×0.100.055
    direct0.00×0.200.000

    2 matching tag(s)

  10. 0.45LessWrong21hnew
    Different fine-tuning objectives install different (and differently fragile) refusal circuits

    This post is a condensed version of our EMNLP 2026 paper - you can take a look here: https://arxiv.org/abs/2609.03887 TL;DR Claim: Different post-training methods install different refusal circuits, and they are…

    why score 0.452
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.74×0.250.185
    contributability0.02×0.150.003
    venue0.39×0.100.039
    direct0.00×0.200.000

    tier-2: circuit; 1 matching tag(s)

  11. 0.40LessWrong13d
    Your Evaluation's Fake names Should Be Unclaimable ,Not Merely Used

    A lot of the talk about the cyber evaluations of July and August revolves around model beliefs and rationalizations. I would like to focus on the harness portion here, especially on one remedy that fixes the wrong…

    why score 0.404
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.003
    contributability0.27×0.150.040
    venue0.61×0.100.061
    direct0.00×0.200.000

    3 matching tag(s)

  12. 0.39Hacker News6d
    Natural emergent misalignment from reward hacking
    why score 0.394
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.11×0.250.028
    contributability0.02×0.150.003
    venue0.13×0.100.013
    direct1.00×0.200.200

    tier-1: reward hacking

  13. 0.39Hacker News7d
    Can escalation channels redirect reward hacking toward defect disclosure?
    why score 0.394
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.08×0.250.021
    contributability0.02×0.150.003
    venue0.21×0.100.021
    direct1.00×0.200.200

    tier-1: reward hacking

  14. 0.38LessWrong7d
    When Activation Oracles learn not to read: Concept-Specific Blind Spots in Fine-Tuned Oracles

    TL;DR Activation Oracles (AOs) are language models trained to answer natural-language questions about another model’s (with the same architecture) internal activations (Karvonen et al. 2025). This way, activation…

    why score 0.384
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.08×0.250.019
    contributability0.12×0.150.017
    venue0.47×0.100.047
    direct0.00×0.200.000

    tier-2: activation; 2 matching tag(s)

  15. 0.38LessWrong7d
    Towards deployment-time misalignment continuation evals: lessons from recent loss of control incidents

    In this post, I will extend the conceptual and methodological work done previously on deployment-time misalignment continuation evals, and lay out a framework for us to think about misalignment spread. I'm currently an…

    why score 0.383
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.09×0.250.022
    contributability0.02×0.150.003
    venue0.58×0.100.058
    direct0.00×0.200.000

    2 matching tag(s)

  16. 0.38LessWrong6hnew
    Training on probes: What's going on

    TL;DR If you train a probe for some property (like "honesty") and do gradient descent against this probe while continuing training that incentivizes dishonesty, the model will change its internal representation to evade…

    why score 0.377
    signalvalueweightpoints
    topic0.25×0.300.075
    liveness0.91×0.250.227
    contributability0.02×0.150.003
    venue0.73×0.100.073
    direct0.00×0.200.000

    tier-2: probe

  17. 0.38LessWrong6d
    BLOOM-WILT: on-policy examples of any LLM behaviour, from a one-line description and logits alone

    Summary of my AI safety paper, BLOOM-WILT, written under the supervision of Edoardo Manino. Code is on GitHub, all transcripts are on HuggingFace, and LogitTilt is now available as part of Inspect. Reach out if you want…

    why score 0.375
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.13×0.250.033
    contributability0.02×0.150.003
    venue0.39×0.100.039
    direct0.00×0.200.000

    2 matching tag(s)

  18. 0.38LessWrong7d
    You can rarely pet the dog in an LLM-generated game

    I asked 17 LLMs to generate browser games that feature a dog. Across 804 playable games generated, the models almost never made the background dog interactive unless they were hinted to do so. I built DogLM, a benchmark…

    why score 0.375
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.09×0.250.023
    contributability0.02×0.150.003
    venue0.50×0.100.050
    direct0.00×0.200.000

    tier-2: feature; 2 matching tag(s)

  19. 0.37Hacker News2w
    Mitigating Reward Hacking as Institutional Design
    why score 0.374
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.000
    contributability0.02×0.150.003
    venue0.21×0.100.021
    direct1.00×0.200.200

    tier-1: reward hacking

  20. 0.37LessWrong2dnew
    Peer Preservation in LLMs: A Replication And Deep Dive

    This work was done as part of the Second Look Fellowship and mentored by Uzay Macar. I'm immensely grateful for the multiple rounds of feedback and support given by Yixiong and Zephaniah Roe for my work. I'm also very…

    why score 0.369
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.43×0.250.108
    contributability0.27×0.150.040
    venue0.72×0.100.072
    direct0.00×0.200.000

    1 matching tag(s)

  21. 0.37LessWrong12d
    Anchoring one concept in a transformer

    We anchored a single concept (red) in the residual stream of a transformer. It ended up where we wanted, with nearby colors graded sensibly, and without degrading task accuracy. Steering is next. Earlier posts in this…

    why score 0.367
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.004
    contributability0.02×0.150.003
    venue0.61×0.100.061
    direct0.00×0.200.000

    2 matching tag(s)

  22. 0.36LessWrong8d
    TASTE: Can AI Models Judge AI Safety Research Proposals?

    tl;dr We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers.…

    why score 0.364
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.06×0.250.015
    contributability0.86×0.150.130
    venue0.70×0.100.070
    direct0.00×0.200.000

    1 matching tag(s)

  23. 0.36LessWrong2w
    LLMs could control their host machines by exploiting inference engines

    Large language models often take actions running on one computer (via an agentic harness such as Claude Code or Codex), however the LLMs’ responses to prompts are computed on a different computer with GPU access. Could…

    why score 0.363
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.01×0.250.002
    contributability0.86×0.150.130
    venue0.81×0.100.081
    direct0.00×0.200.000

    1 matching tag(s)

  24. 0.35LessWrong2w
    Rogue Scalpel: Activation steering breaks refusal, even with benign directions

    TLDR: Activation steering can bypass refusal even when the steering vector represents a benign concept like "brand identity" or "Portugal". It is hard (maybe impossible) to predict which vector will bypass refusal on…

    why score 0.353
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.02×0.150.003
    venue0.50×0.100.050
    direct0.00×0.200.000

    tier-2: feature, activation; 1 matching tag(s)

  25. 0.34LessWrong13d
    Your Agent's Trace Probably Cannot Tell You Who Approved a Tool Call

    Consider some agent stack that you maintain. Find a tool call within its trace log. Can you determine if a person authorized it or if some policy automatically waived that requirement? All classifiers over-block in…

    why score 0.342
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.003
    contributability0.02×0.150.003
    venue0.37×0.100.037
    direct0.00×0.200.000

    3 matching tag(s)