Field Radar.

What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.

How this list is made

This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.

The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.

Sources — LessWrong: error: HTTP 429 for https://www.lesswrong.com/graphql · Hacker News: ok (5 stories) · Reddit: ok (25 posts) — some subreddits rate-limited

As of
2026-08-08 14:00 ET
Showing
25 items
New
0 in last 48h
Refresh
Every 6 hours
  1. 0.62LessWrong4d
    RLVR that rewards red teaming the training environment

    Epistemic status: throwing an idea at the wall and seeing if it sticks I've been thinking about how to mitigate egregious reward hacking, a la the Hugging Face incident. I don't have the resources I'd need to write a…

    why score 0.616
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.21×0.250.052
    contributability0.86×0.150.130
    venue0.85×0.100.085
    direct1.00×0.200.200

    tier-1: reward hacking

  2. 0.55LessWrong2w
    A Multi-Agent Extension for Petri

    Intro Petri is an open-source framework built on Inspect AI for automated AI Safety evaluations first released by Anthropic, but now maintained and developed by Meridian Labs. Each evaluation involves three agents, the…

    why score 0.549
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.02×0.150.003
    venue0.45×0.100.045
    direct1.00×0.200.200

    2 matching tag(s)

  3. 0.42Hacker News5d
    The Download: reward hacking explained, and suspected Iranian cyberattacks
    why score 0.420
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.19×0.250.047
    contributability0.02×0.150.003
    venue0.21×0.100.021
    direct1.00×0.200.200

    tier-1: reward hacking

  4. 0.42LessWrong4d
    When should we trust a latent representation?

    Working on AI Safety, I spend a significant part of my time thinking about how frontier AI Safety research eventually gets translated into deployable safety systems. One thing I have constantly noticed is the transition…

    why score 0.419
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.24×0.250.060
    contributability0.12×0.150.017
    venue0.41×0.100.041
    direct0.00×0.200.000

    2 matching tag(s)

  5. 0.42LessWrong4d
    Rewrite All the Code, All the Time

    This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. It has become a mainstream prediction that…

    why score 0.418
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.26×0.250.065
    contributability1.00×0.150.149
    venue0.53×0.100.053
    direct0.00×0.200.000

    1 matching tag(s)

  6. 0.42LessWrong2w
    Fable is SOTA at CIFAR Speedrun (& specification gaming)

    Fulcrum is working on an AI R&D optimization benchmark. Here, we present results from one of our tasks, including preliminary results from Fable. For more detail on Fable’s solution, check out…

    why score 0.417
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.000
    contributability0.02×0.150.003
    venue0.64×0.100.064
    direct1.00×0.200.200

    tier-1: specification gaming

  7. 0.41LessWrong5d
    Single Forward Pass Evals on Fable, Opus 5, and GPT-5.6-Sol

    This is a research update for an on-going replication of single-forward-pass evals done as part of the Second Look Fellowship. In following posts, we will run more comprehensive replications of previous work and release…

    why score 0.407
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.14×0.250.036
    contributability0.02×0.150.003
    venue0.69×0.100.069
    direct0.00×0.200.000

    2 matching tag(s)

  8. 0.40LessWrong4d
    Deliberate Alignment Faking as a Defense Against Model Poisoning

    I want to discuss and brainstorm a counterintuitive approach to AI alignment: Inducing alignment faking on purpose, to prevent the model from developing emergent misalignment. To prevent this from going horribly wrong,…

    why score 0.401
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.19×0.250.049
    contributability0.02×0.150.003
    venue0.50×0.100.050
    direct0.00×0.200.000

    2 matching tag(s)

  9. 0.40LessWrong11d
    Multi-Turn Drift Increases Scheming

    TLDR - We talk about scheming, and why research on this phenomenon is crucial for AI safety. We find a particular environment/scenario where scheming happens at a higher rate than normal. We provide hypotheses for why…

    why score 0.395
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.02×0.250.005
    contributability0.27×0.150.040
    venue0.51×0.100.051
    direct0.00×0.200.000

    2 matching tag(s)

  10. 0.38Hacker News2w
    Show HN: Vinv-Ties every runtime trace to code segment, prevents reward hacking
    why score 0.380
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.01×0.250.002
    contributability0.02×0.150.003
    venue0.26×0.100.026
    direct1.00×0.200.200

    tier-1: reward hacking

  11. 0.38Hacker News2w
    Show HN: VinvAI – Ties runtime trace to code segment, prevents reward hacking
    why score 0.375
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.01×0.250.002
    contributability0.02×0.150.003
    venue0.21×0.100.021
    direct1.00×0.200.200

    tier-1: reward hacking

  12. 0.37LessWrong2w
    Linear probes tell you where quantization will hurt

    Epistemic status: I have only tested one encoder family (BERT-base and its relatives) and one decoder LLM (Qwen2.5-3B), one seed, token-level tasks, and post-training weight quantization. I trust the results because I…

    why score 0.367
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.002
    contributability0.02×0.150.003
    venue0.63×0.100.063
    direct0.00×0.200.000

    2 matching tag(s)

  13. 0.36LessWrong4d
    Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face

    Three-Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we describe the ambitious,…

    why score 0.364
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.24×0.250.060
    contributability0.42×0.150.064
    venue0.90×0.100.090
    direct0.00×0.200.000

    1 matching tag(s)

  14. 0.36LessWrong2w
    V&V takes on OpenAI’s long-horizon incidents

    [Cross-posted from The Foretellix CTO Blog. These short takes try to put a verification-and-validation slant on AI-safety / alignment topics – they are not full treatments. I co-originated coverage-driven verification…

    why score 0.363
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.02×0.150.003
    venue0.59×0.100.059
    direct0.00×0.200.000

    2 matching tag(s)

  15. 0.36LessWrong2w
    Is there even a ground-truth for LLMs’ internal representations?

    [This is an introductory blog for the paper Laguerre Geometry for Interpreting Large Language Models and the GitHub repository Geometric Lens.] LLM Lens: What does an internal vector mean? Anthropic's recent paper on…

    why score 0.363
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.000
    contributability0.02×0.150.003
    venue0.60×0.100.060
    direct0.00×0.200.000

    2 matching tag(s)

  16. 0.36LessWrong2w
    The State of AI Consciousness Research

    Epistemic status: a survey, not an argument. I am agnostic on whether any current system is conscious; the claim is only that the question is researchable. This piece surveys the empirical research on AI consciousness.…

    why score 0.362
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.00×0.250.000
    contributability0.42×0.150.064
    venue0.73×0.100.073
    direct0.00×0.200.000

    tier-2: mechanistic interpretability; 1 matching tag(s)

  17. 0.36LessWrong2w
    SONI: Selective Orthogonalisation via Noise Injection

    This project was completed as a capstone for TARA. All code is available in github. TL;DR The Problem: Neural networks use superposition to pack many concepts into small latent spaces by making feature vectors…

    why score 0.362
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.002
    contributability0.02×0.150.003
    venue0.57×0.100.057
    direct0.00×0.200.000

    tier-2: feature, activation; 1 matching tag(s)

  18. 0.36r/ControlProblem11d
    What failed in the Hugging Face incident wasn't the model

    TL/DR: everyone is arguing about whether the Hugging Face models went rogue. They didn't, it's specification gaming, that argument is a decade old. The part being missed is that three separate checks sat above those…

    why score 0.359
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.02×0.250.006
    contributability0.02×0.150.003
    venue0.00×0.100.000
    direct1.00×0.200.200

    tier-1: specification gaming

  19. 0.36LessWrong2w
    Mechanistic interpretability hypotheses for Measuring Reward-Seeking by Instilling Contrastive Beliefs and additional comments

    This is interesting research! https://alignment.openai.com/measuring-reward-seeking It made me think of few overlapping hypotheses for what might be happening here, how did the grader behavior emerge at the pretraining…

    why score 0.357
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.02×0.150.003
    venue0.53×0.100.053
    direct0.00×0.200.000

    tier-2: mechanistic interpretability; 2 matching tag(s)

  20. 0.36r/ControlProblem13d
    Is specification downstream from judgment?

    Hey everyone. I’ve long been fascinated by both philosophy of technology and AI alignment. I’m also using Heidegger quite a bit for my philosophy PhD. Given the recent OpenAI–Hugging Face incident reported this week, I…

    why score 0.355
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.01×0.250.003
    contributability0.02×0.150.003
    venue0.00×0.100.000
    direct1.00×0.200.200

    tier-1: specification gaming

  21. 0.36LessWrong12d
    LLMs are (still) mostly powered by imitative learning, not RL

    Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It’s so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture. Stepping back,…

    why score 0.355
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.01×0.250.003
    contributability0.72×0.150.108
    venue0.94×0.100.094
    direct0.00×0.200.000

    1 matching tag(s)

  22. 0.35LessWrong12d
    Inoculate or Reflect? Two training interventions under prompting, steering, and patching

    Anthropic's recent paper, Verbalizable Representations Form a Global Workspace in Language Models, contains a small experiment near the end that we found more interesting than the main findings. Surprising that it's so…

    why score 0.349
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.003
    contributability0.02×0.150.003
    venue0.43×0.100.043
    direct0.00×0.200.000

    2 matching tag(s)

  23. 0.35LessWrong2w
    Steering Blackmail Through a Model's "Emotional State"

    Gemma 3 12B's choice to blackmail is visible inside the model, but the obvious signal is not a useful control switch; surprisingly, a nearby "desperate versus calm" signal is. In this case study, I show how that…

    why score 0.348
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.02×0.150.003
    venue0.45×0.100.045
    direct0.00×0.200.000

    2 matching tag(s)

  24. 0.35LessWrong2w
    Tracing causal structure in LLM-generated text: a different lens on the Dallas circuit

    The classic "Dallas" example from Anthropic focuses on an internal circuit in an LLM. I became curious about what the same underlying process looks like when viewed through the generated reasoning trace instead of…

    why score 0.348
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.000
    contributability0.02×0.150.003
    venue0.45×0.100.045
    direct0.00×0.200.000

    tier-2: circuit; 2 matching tag(s)

  25. 0.35LessWrong11d
    When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

    By Sai Kartheek Reddy Kasu, Nils Lukas, and Samuele Poppi This post is a summary of our accepted paper at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN). The full paper is available here TL;DR The Setup:…

    why score 0.346
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.02×0.250.005
    contributability0.02×0.150.003
    venue0.39×0.100.039
    direct0.00×0.200.000

    2 matching tag(s)