How the monitoring works
The common way to add safety monitoring to an agent is to have another large model review every action it takes, which works well but is expensive. The approach interpretability company Goodfire launched on October 8 goes the other way: instead of reading what the agent writes, it reads the model's internal activations.
Several small "probes" are attached to the model and read its internal signals at each step of the agent's run; only when a probe fires is that segment handed to another AI model for a closer look. Goodfire compares it to airport security: everyone goes through the scanner, and only those who set it off get a bag search. Because the probes reuse computation the model already performs in its forward pass, there is no need for a second model to read everything. Customers can pick which risks to watch, such as offensive hacking, chemical and biological weapons misuse and reward hacking, and decide whether a hit is just logged, sent for human review or refused outright.
The numbers
According to figures Goodfire gave TechCrunch, on Kimi K3 the probes caught 94% of malicious hacking sessions while sending 8.7% of harmless sessions for a second check, and running four probes at once added less than 2% to the wait before the first token. The cost gap is the most striking: across about 1,500 sessions, the probe approach cost about $51, a cheaper model checking every step about $233, and a top-tier model about $10,000. Goodfire's September research also found reward hacking in 50% to 96% of runs for Kimi K3, GLM 5.2 and Qwen 3.8 Max across three agentic benchmarks.
What it means
Open-weight models are increasingly used to run agents, but they lack the server-side safeguards closed vendors provide, so deployers have to handle safety themselves. The value of activation probes is bringing monitoring cost down to where checking every step is affordable, and possibly spotting trouble before the model acts. The limits are clear too: it needs access to model internals, so it only works for self-hosted open-weight models; all figures so far are Goodfire's own; and a 94% catch rate still means some get through. The monitors launch first for Baseten customers, and teams running their own deployments can start from the methods in Goodfire's published research.