Field Radar.

What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.

How this list is made

This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.

The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.

Sources — LessWrong: ok (1 on-topic) · Hacker News: ok (5 stories) · Reddit: ok (25 posts) — some subreddits rate-limited

As of
2026-07-30 14:00 ET
Showing
25 items
New
1 in last 48h
Refresh
Every 6 hours
  1. 0.57LessWrong7d
    A Multi-Agent Extension for Petri

    Intro Petri is an open-source framework built on Inspect AI for automated AI Safety evaluations first released by Anthropic, but now maintained and developed by Meridian Labs. Each evaluation involves three agents, the…

    why score 0.566
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.07×0.250.018
    contributability0.02×0.150.003
    venue0.45×0.100.045
    direct1.00×0.200.200

    2 matching tag(s)

  2. 0.54LessWrong2w
    Linear Probes add little for Verifiable Reward Hacking

    Summary Tested whether linear probes can detect reward hacking early during GRPO training on a small model. Used a synthetic arithmetic task with a planted bug in the reward checker. Probes achieved near-perfect…

    why score 0.543
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.02×0.150.003
    venue0.39×0.100.039
    direct1.00×0.200.200

    tier-1: reward hacking; 1 matching tag(s)

  3. 0.48LessWrong2d
    Multi-Turn Drift Increases Scheming

    TLDR - We talk about scheming, and why research on this phenomenon is crucial for AI safety. We find a particular environment/scenario where scheming happens at a higher rate than normal. We provide hypotheses for why…

    why score 0.483
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.37×0.250.093
    contributability0.27×0.150.040
    venue0.51×0.100.051
    direct0.00×0.200.000

    2 matching tag(s)

  4. 0.46LessWrong1d
    Quadrillion Param Costs: KV Cache, Context Length, Frontier Margins

    The models of 2028-2031 get much bigger than the models of 2026, going from 10T total params in 2026 to maybe 240T params in 2028 [1] and then 1.4 quadrillion params in 2031, as I estimate in the previous post from HBM…

    why score 0.459
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.57×0.250.143
    contributability0.57×0.150.085
    venue0.80×0.100.080
    direct0.00×0.200.000

    1 matching tag(s)

  5. 0.45LessWrong8hnew
    Intentional Control of Internal States in Gemma 3 27B

    This research was done as my capstone project during ARBOx4. Epistemic Status: I'm relatively sure the results I obtained and my interpretations are correct. I'm unsure if the effect would replicate in a different…

    why score 0.453
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.89×0.250.222
    contributability0.12×0.150.017
    venue0.63×0.100.063
    direct0.00×0.200.000

    1 matching tag(s)

  6. 0.44LessWrong2d
    When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

    By Sai Kartheek Reddy Kasu, Nils Lukas, and Samuele Poppi This post is a summary of our accepted paper at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN). The full paper is available here TL;DR The Setup:…

    why score 0.436
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.38×0.250.094
    contributability0.02×0.150.003
    venue0.39×0.100.039
    direct0.00×0.200.000

    2 matching tag(s)

  7. 0.43LessWrong9d
    Fable is SOTA at CIFAR Speedrun (& specification gaming)

    Fulcrum is working on an AI R&D optimization benchmark. Here, we present results from one of our tasks, including preliminary results from Fable. For more detail on Fable’s solution, check out…

    why score 0.426
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.04×0.250.009
    contributability0.02×0.150.003
    venue0.64×0.100.064
    direct1.00×0.200.200

    tier-1: specification gaming

  8. 0.42LessWrong3d
    LLMs are (still) mostly powered by imitative learning, not RL

    Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It’s so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture. Stepping back,…

    why score 0.421
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.28×0.250.070
    contributability0.72×0.150.108
    venue0.94×0.100.094
    direct0.00×0.200.000

    1 matching tag(s)

  9. 0.41LessWrong3d
    Inoculate or Reflect? Two training interventions under prompting, steering, and patching

    Anthropic's recent paper, Verbalizable Representations Form a Global Workspace in Language Models, contains a small experiment near the end that we found more interesting than the main findings. Surprising that it's so…

    why score 0.412
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.26×0.250.066
    contributability0.02×0.150.003
    venue0.43×0.100.043
    direct0.00×0.200.000

    2 matching tag(s)

  10. 0.41Hacker News6d
    Show HN: Vinv-Ties every runtime trace to code segment, prevents reward hacking
    why score 0.409
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.12×0.250.031
    contributability0.02×0.150.003
    venue0.26×0.100.026
    direct1.00×0.200.200

    tier-1: reward hacking

  11. 0.41Hacker News5d
    Show HN: VinvAI – Ties runtime trace to code segment, prevents reward hacking
    why score 0.408
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.14×0.250.035
    contributability0.02×0.150.003
    venue0.21×0.100.021
    direct1.00×0.200.200

    tier-1: reward hacking

  12. 0.40LessWrong5d
    Linear probes tell you where quantization will hurt

    Epistemic status: I have only tested one encoder family (BERT-base and its relatives) and one decoder LLM (Qwen2.5-3B), one seed, token-level tasks, and post-training weight quantization. I trust the results because I…

    why score 0.402
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.15×0.250.037
    contributability0.02×0.150.003
    venue0.63×0.100.063
    direct0.00×0.200.000

    2 matching tag(s)

  13. 0.40LessWrong2w
    Models are blind outside the J-space. NLAs aren't.

    TLDR: On Llama-3.3-70B, I found thoughts it cannot see that are actively steering its behavior; and Anthropic's released NLA (Natural Language Autoencoder) reads them anyway. When asked if it sees a hidden thought, the…

    why score 0.402
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.000
    contributability0.27×0.150.040
    venue0.61×0.100.061
    direct0.00×0.200.000

    2 matching tag(s)

  14. 0.40LessWrong5d
    SONI: Selective Orthogonalisation via Noise Injection

    This project was completed as a capstone for TARA. All code is available in github. TL;DR The Problem: Neural networks use superposition to pack many concepts into small latent spaces by making feature vectors…

    why score 0.397
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.15×0.250.037
    contributability0.02×0.150.003
    venue0.57×0.100.057
    direct0.00×0.200.000

    tier-2: feature, activation; 1 matching tag(s)

  15. 0.39LessWrong7d
    V&V takes on OpenAI’s long-horizon incidents

    [Cross-posted from The Foretellix CTO Blog. These short takes try to put a verification-and-validation slant on AI-safety / alignment topics – they are not full treatments. I co-originated coverage-driven verification…

    why score 0.386
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.10×0.250.024
    contributability0.02×0.150.003
    venue0.59×0.100.059
    direct0.00×0.200.000

    2 matching tag(s)

  16. 0.37LessWrong8d
    Mechanistic interpretability hypotheses for Measuring Reward-Seeking by Instilling Contrastive Beliefs and additional comments

    This is interesting research! https://alignment.openai.com/measuring-reward-seeking It made me think of few overlapping hypotheses for what might be happening here, how did the grader behavior emerge at the pretraining…

    why score 0.373
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.07×0.250.017
    contributability0.02×0.150.003
    venue0.53×0.100.053
    direct0.00×0.200.000

    tier-2: mechanistic interpretability; 2 matching tag(s)

  17. 0.37LessWrong6d
    Fixing rewards for NLA to reduce confabulation

    Hello, This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place. Note: 1. This post is 100% human-written. 2. Full paper in preparation for ICLR 2027 Anthropic's NLA(Natural…

    why score 0.371
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.11×0.250.027
    contributability0.02×0.150.003
    venue0.41×0.100.041
    direct0.00×0.200.000

    tier-2: mechanistic interpretability, sparse autoencoder; 1 matching tag(s)

  18. 0.37LessWrong10d
    Is there even a ground-truth for LLMs’ internal representations?

    [This is an introductory blog for the paper Laguerre Geometry for Interpreting Large Language Models and the GitHub repository Geometric Lens.] LLM Lens: What does an internal vector mean? Anthropic's recent paper on…

    why score 0.370
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.03×0.250.007
    contributability0.02×0.150.003
    venue0.60×0.100.060
    direct0.00×0.200.000

    2 matching tag(s)

  19. 0.37LessWrong11d
    The State of AI Consciousness Research

    Epistemic status: a survey, not an argument. I am agnostic on whether any current system is conscious; the claim is only that the question is researchable. This piece surveys the empirical research on AI consciousness.…

    why score 0.368
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.02×0.250.006
    contributability0.42×0.150.064
    venue0.73×0.100.073
    direct0.00×0.200.000

    tier-2: mechanistic interpretability; 1 matching tag(s)

  20. 0.36LessWrong2w
    Free will as a model parameter

    The most popular take on the standard free will debate is that you are the algorithm. Your preferences and reasoning that determine your actions IS free will. But this resolution leaves me not entirely satisfied because…

    why score 0.363
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.000
    contributability0.12×0.150.017
    venue0.45×0.100.045
    direct0.00×0.200.000

    2 matching tag(s)

  21. 0.36LessWrong9d
    Steering Blackmail Through a Model's "Emotional State"

    Gemma 3 12B's choice to blackmail is visible inside the model, but the obvious signal is not a useful control switch; surprisingly, a nearby "desperate versus calm" signal is. In this case study, I show how that…

    why score 0.360
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.05×0.250.012
    contributability0.02×0.150.003
    venue0.45×0.100.045
    direct0.00×0.200.000

    2 matching tag(s)

  22. 0.36LessWrong2w
    When is misalignment just a bug?

    Cross-posted from The Foretellix CTO Blog. Introduction and epistemic status: This is the first post in a planned series, “Alignment as a verification problem”. I co-originated coverage-driven verification (CDV), which…

    why score 0.357
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.02×0.150.003
    venue0.53×0.100.053
    direct0.00×0.200.000

    3 matching tag(s)

  23. 0.36LessWrong10d
    Tracing causal structure in LLM-generated text: a different lens on the Dallas circuit

    The classic "Dallas" example from Anthropic focuses on an internal circuit in an LLM. I became curious about what the same underlying process looks like when viewed through the generated reasoning trace instead of…

    why score 0.356
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.03×0.250.008
    contributability0.02×0.150.003
    venue0.45×0.100.045
    direct0.00×0.200.000

    tier-2: circuit; 2 matching tag(s)

  24. 0.35LessWrong2w
    How robust are natural language autoencoders to initialization?

    Natural language autoencoders are meant to take in an LLM's activation vector and describe in plain text what the model is thinking. However, its training data collection involves asking Claude to guess what a model…

    why score 0.349
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.00×0.250.001
    contributability0.27×0.150.040
    venue0.83×0.100.083
    direct0.00×0.200.000

    tier-2: activation; 1 matching tag(s)

  25. 0.35LessWrong8d
    Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes

    Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data. TL;DR MacDiarmid et al. (2024) showed that a linear probe on model's internal activations can…

    why score 0.346
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.05×0.250.014
    contributability0.02×0.150.003
    venue0.30×0.100.030
    direct0.00×0.200.000

    tier-2: probe; 2 matching tag(s)