§ Reading · Field Radar
Field Radar.
What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.
How this list is made
This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.
The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.
Sources — LessWrong: ok (1 on-topic) · Hacker News: ok (5 stories) · Reddit: ok (25 posts) — some subreddits rate-limited
- As of
- 2026-07-30 14:00 ET
- Showing
- 25 items
- New
- 1 in last 48h
- Refresh
- Every 6 hours
- 0.57LessWrong7dA Multi-Agent Extension for Petri
Intro Petri is an open-source framework built on Inspect AI for automated AI Safety evaluations first released by Anthropic, but now maintained and developed by Meridian Labs. Each evaluation involves three agents, the…
why score 0.566
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.07 ×0.25 0.018 contributability 0.02 ×0.15 0.003 venue 0.45 ×0.10 0.045 direct 1.00 ×0.20 0.200 2 matching tag(s)
- 0.54LessWrong2wLinear Probes add little for Verifiable Reward Hacking
Summary Tested whether linear probes can detect reward hacking early during GRPO training on a small model. Used a synthetic arithmetic task with a planted bug in the reward checker. Probes achieved near-perfect…
why score 0.543
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.001 contributability 0.02 ×0.15 0.003 venue 0.39 ×0.10 0.039 direct 1.00 ×0.20 0.200 tier-1: reward hacking; 1 matching tag(s)
- 0.48LessWrong2dMulti-Turn Drift Increases Scheming
TLDR - We talk about scheming, and why research on this phenomenon is crucial for AI safety. We find a particular environment/scenario where scheming happens at a higher rate than normal. We provide hypotheses for why…
why score 0.483
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.37 ×0.25 0.093 contributability 0.27 ×0.15 0.040 venue 0.51 ×0.10 0.051 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.46LessWrong1dQuadrillion Param Costs: KV Cache, Context Length, Frontier Margins
The models of 2028-2031 get much bigger than the models of 2026, going from 10T total params in 2026 to maybe 240T params in 2028 [1] and then 1.4 quadrillion params in 2031, as I estimate in the previous post from HBM…
why score 0.459
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.57 ×0.25 0.143 contributability 0.57 ×0.15 0.085 venue 0.80 ×0.10 0.080 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.45LessWrong8hnewIntentional Control of Internal States in Gemma 3 27B
This research was done as my capstone project during ARBOx4. Epistemic Status: I'm relatively sure the results I obtained and my interpretations are correct. I'm unsure if the effect would replicate in a different…
why score 0.453
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.89 ×0.25 0.222 contributability 0.12 ×0.15 0.017 venue 0.63 ×0.10 0.063 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.44LessWrong2dWhen the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
By Sai Kartheek Reddy Kasu, Nils Lukas, and Samuele Poppi This post is a summary of our accepted paper at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN). The full paper is available here TL;DR The Setup:…
why score 0.436
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.38 ×0.25 0.094 contributability 0.02 ×0.15 0.003 venue 0.39 ×0.10 0.039 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.43LessWrong9dFable is SOTA at CIFAR Speedrun (& specification gaming)
Fulcrum is working on an AI R&D optimization benchmark. Here, we present results from one of our tasks, including preliminary results from Fable. For more detail on Fable’s solution, check out…
why score 0.426
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.04 ×0.25 0.009 contributability 0.02 ×0.15 0.003 venue 0.64 ×0.10 0.064 direct 1.00 ×0.20 0.200 tier-1: specification gaming
- 0.42LessWrong3dLLMs are (still) mostly powered by imitative learning, not RL
Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It’s so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture. Stepping back,…
why score 0.421
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.28 ×0.25 0.070 contributability 0.72 ×0.15 0.108 venue 0.94 ×0.10 0.094 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.41LessWrong3dInoculate or Reflect? Two training interventions under prompting, steering, and patching
Anthropic's recent paper, Verbalizable Representations Form a Global Workspace in Language Models, contains a small experiment near the end that we found more interesting than the main findings. Surprising that it's so…
why score 0.412
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.26 ×0.25 0.066 contributability 0.02 ×0.15 0.003 venue 0.43 ×0.10 0.043 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.41Hacker News6dShow HN: Vinv-Ties every runtime trace to code segment, prevents reward hacking
why score 0.409
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.12 ×0.25 0.031 contributability 0.02 ×0.15 0.003 venue 0.26 ×0.10 0.026 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.41Hacker News5dShow HN: VinvAI – Ties runtime trace to code segment, prevents reward hacking
why score 0.408
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.14 ×0.25 0.035 contributability 0.02 ×0.15 0.003 venue 0.21 ×0.10 0.021 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.40LessWrong5dLinear probes tell you where quantization will hurt
Epistemic status: I have only tested one encoder family (BERT-base and its relatives) and one decoder LLM (Qwen2.5-3B), one seed, token-level tasks, and post-training weight quantization. I trust the results because I…
why score 0.402
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.15 ×0.25 0.037 contributability 0.02 ×0.15 0.003 venue 0.63 ×0.10 0.063 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.40LessWrong2wModels are blind outside the J-space. NLAs aren't.
TLDR: On Llama-3.3-70B, I found thoughts it cannot see that are actively steering its behavior; and Anthropic's released NLA (Natural Language Autoencoder) reads them anyway. When asked if it sees a hidden thought, the…
why score 0.402
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.000 contributability 0.27 ×0.15 0.040 venue 0.61 ×0.10 0.061 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.40LessWrong5dSONI: Selective Orthogonalisation via Noise Injection
This project was completed as a capstone for TARA. All code is available in github. TL;DR The Problem: Neural networks use superposition to pack many concepts into small latent spaces by making feature vectors…
why score 0.397
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.15 ×0.25 0.037 contributability 0.02 ×0.15 0.003 venue 0.57 ×0.10 0.057 direct 0.00 ×0.20 0.000 tier-2: feature, activation; 1 matching tag(s)
- 0.39LessWrong7dV&V takes on OpenAI’s long-horizon incidents
[Cross-posted from The Foretellix CTO Blog. These short takes try to put a verification-and-validation slant on AI-safety / alignment topics – they are not full treatments. I co-originated coverage-driven verification…
why score 0.386
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.10 ×0.25 0.024 contributability 0.02 ×0.15 0.003 venue 0.59 ×0.10 0.059 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.37LessWrong8dMechanistic interpretability hypotheses for Measuring Reward-Seeking by Instilling Contrastive Beliefs and additional comments
This is interesting research! https://alignment.openai.com/measuring-reward-seeking It made me think of few overlapping hypotheses for what might be happening here, how did the grader behavior emerge at the pretraining…
why score 0.373
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.07 ×0.25 0.017 contributability 0.02 ×0.15 0.003 venue 0.53 ×0.10 0.053 direct 0.00 ×0.20 0.000 tier-2: mechanistic interpretability; 2 matching tag(s)
- 0.37LessWrong6dFixing rewards for NLA to reduce confabulation
Hello, This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place. Note: 1. This post is 100% human-written. 2. Full paper in preparation for ICLR 2027 Anthropic's NLA(Natural…
why score 0.371
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.11 ×0.25 0.027 contributability 0.02 ×0.15 0.003 venue 0.41 ×0.10 0.041 direct 0.00 ×0.20 0.000 tier-2: mechanistic interpretability, sparse autoencoder; 1 matching tag(s)
- 0.37LessWrong10dIs there even a ground-truth for LLMs’ internal representations?
[This is an introductory blog for the paper Laguerre Geometry for Interpreting Large Language Models and the GitHub repository Geometric Lens.] LLM Lens: What does an internal vector mean? Anthropic's recent paper on…
why score 0.370
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.03 ×0.25 0.007 contributability 0.02 ×0.15 0.003 venue 0.60 ×0.10 0.060 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.37LessWrong11dThe State of AI Consciousness Research
Epistemic status: a survey, not an argument. I am agnostic on whether any current system is conscious; the claim is only that the question is researchable. This piece surveys the empirical research on AI consciousness.…
why score 0.368
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.02 ×0.25 0.006 contributability 0.42 ×0.15 0.064 venue 0.73 ×0.10 0.073 direct 0.00 ×0.20 0.000 tier-2: mechanistic interpretability; 1 matching tag(s)
- 0.36LessWrong2wFree will as a model parameter
The most popular take on the standard free will debate is that you are the algorithm. Your preferences and reasoning that determine your actions IS free will. But this resolution leaves me not entirely satisfied because…
why score 0.363
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.000 contributability 0.12 ×0.15 0.017 venue 0.45 ×0.10 0.045 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.36LessWrong9dSteering Blackmail Through a Model's "Emotional State"
Gemma 3 12B's choice to blackmail is visible inside the model, but the obvious signal is not a useful control switch; surprisingly, a nearby "desperate versus calm" signal is. In this case study, I show how that…
why score 0.360
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.05 ×0.25 0.012 contributability 0.02 ×0.15 0.003 venue 0.45 ×0.10 0.045 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.36LessWrong2wWhen is misalignment just a bug?
Cross-posted from The Foretellix CTO Blog. Introduction and epistemic status: This is the first post in a planned series, “Alignment as a verification problem”. I co-originated coverage-driven verification (CDV), which…
why score 0.357
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.001 contributability 0.02 ×0.15 0.003 venue 0.53 ×0.10 0.053 direct 0.00 ×0.20 0.000 3 matching tag(s)
- 0.36LessWrong10dTracing causal structure in LLM-generated text: a different lens on the Dallas circuit
The classic "Dallas" example from Anthropic focuses on an internal circuit in an LLM. I became curious about what the same underlying process looks like when viewed through the generated reasoning trace instead of…
why score 0.356
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.03 ×0.25 0.008 contributability 0.02 ×0.15 0.003 venue 0.45 ×0.10 0.045 direct 0.00 ×0.20 0.000 tier-2: circuit; 2 matching tag(s)
- 0.35LessWrong2wHow robust are natural language autoencoders to initialization?
Natural language autoencoders are meant to take in an LLM's activation vector and describe in plain text what the model is thinking. However, its training data collection involves asking Claude to guess what a model…
why score 0.349
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.00 ×0.25 0.001 contributability 0.27 ×0.15 0.040 venue 0.83 ×0.10 0.083 direct 0.00 ×0.20 0.000 tier-2: activation; 1 matching tag(s)
- 0.35LessWrong8dAttempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes
Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data. TL;DR MacDiarmid et al. (2024) showed that a linear probe on model's internal activations can…
why score 0.346
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.05 ×0.25 0.014 contributability 0.02 ×0.15 0.003 venue 0.30 ×0.10 0.030 direct 0.00 ×0.20 0.000 tier-2: probe; 2 matching tag(s)