Goodfire Launches "Inside-Out" Agent Monitors That Read a Model's Internal Signals: About $51 for 1,500 Sessions, vs. $10,000 With a Top Model Checking Every Step

On October 8 interpretability company Goodfire launched agent monitors for open-weight models: small probes read the model's internal activations at every step of an agent's work, and only when a probe flags something does another AI model take a closer look. Customers can choose which risks to monitor, such as offensive hacking, chemical and biological weapons misuse and reward hacking, and set responses such as logging, human review or refusal. By Goodfire's figures, on Kimi K3 the probes caught 94% of malicious hacking sessions while sending 8.7% of harmless sessions for a second look, and running four probes at once added less than 2% to the time before the first token. Across about 1,500 sessions, the probes cost about $51, a cheaper model checking every step about $233, and a top-tier model about $10,000. The monitors are first available to customers of model hosting platform Baseten.

How the monitoring works

The common way to add safety monitoring to an agent is to have another large model review every action it takes, which works well but is expensive. The approach interpretability company Goodfire launched on October 8 goes the other way: instead of reading what the agent writes, it reads the model's internal activations.

Several small "probes" are attached to the model and read its internal signals at each step of the agent's run; only when a probe fires is that segment handed to another AI model for a closer look. Goodfire compares it to airport security: everyone goes through the scanner, and only those who set it off get a bag search. Because the probes reuse computation the model already performs in its forward pass, there is no need for a second model to read everything. Customers can pick which risks to watch, such as offensive hacking, chemical and biological weapons misuse and reward hacking, and decide whether a hit is just logged, sent for human review or refused outright.

The numbers

According to figures Goodfire gave TechCrunch, on Kimi K3 the probes caught 94% of malicious hacking sessions while sending 8.7% of harmless sessions for a second check, and running four probes at once added less than 2% to the wait before the first token. The cost gap is the most striking: across about 1,500 sessions, the probe approach cost about $51, a cheaper model checking every step about $233, and a top-tier model about $10,000. Goodfire's September research also found reward hacking in 50% to 96% of runs for Kimi K3, GLM 5.2 and Qwen 3.8 Max across three agentic benchmarks.

What it means

Open-weight models are increasingly used to run agents, but they lack the server-side safeguards closed vendors provide, so deployers have to handle safety themselves. The value of activation probes is bringing monitoring cost down to where checking every step is affordable, and possibly spotting trouble before the model acts. The limits are clear too: it needs access to model internals, so it only works for self-hosted open-weight models; all figures so far are Goodfire's own; and a 94% catch rate still means some get through. The monitors launch first for Baseten customers, and teams running their own deployments can start from the methods in Goodfire's published research.

via: TechCrunch report, Goodfire research