MIT, Sakana AI cut coding-agent eval costs with SIFT framework
MIT and Sakana AI Release Cost-Cutting Evaluation Framework for Coding Agents
Researchers at the Massachusetts Institute of Technology and Sakana AI have published Self-Improvement via Fast Tree-Search (SIFT), a framework designed to slash evaluation costs for self-improving coding agents without sacrificing performance. The approach uses a separate language model to compare candidate agents before expensive benchmark evaluation, according to VentureBeat.
The framework addresses a critical bottleneck in agent development: running full benchmark suites on every candidate agent is computationally expensive and time-consuming. SIFT intercepts this process by filtering candidates through a lightweight LLM judge, which reduces the number of expensive evaluations needed to identify the best-performing agents.
Benchmark Results and Cost Figures
In testing, SIFT achieved 35.1% on Polyglot and 36.7% on TerminalBench 2.1, with reported performance on SWE-bench Verified subset configurations, per AI Weekly. The cost reductions are substantial: one Polyglot evaluation run consumed approximately $150 in API credits and 42 CPU hours. A second configuration used $34 of API spend and 224 CPU hours, while a smaller 50-task Polyglot subset required only $6 and 2.6 CPU hours to evaluate, according to Ground News.
These figures matter because agent development has historically required teams to run dozens or hundreds of candidate agents through full benchmarks—a process that can cost thousands of dollars and consume days of compute time. SIFT's pre-filtering step allows teams to prune weak candidates before the expensive runs, compressing timelines and budgets.
Developer Impact and Adoption Path
The framework is aimed at teams building agents that use iterative self-improvement loops, where agents generate candidate solutions and need rapid feedback on which ones are worth full evaluation. By reducing the cost per evaluation cycle, SIFT enables faster iteration and experimentation—critical for competitive development in the agent economy.
The work is positioned within the broader researcher focus on making agent training and evaluation more efficient. As coding agents become standard infrastructure in enterprise and open-source stacks, cost-effective evaluation becomes a business differentiator. MIT and Sakana AI's contribution is a concrete, measurable tool that developers can adopt immediately, not a theoretical proposal or roadmap announcement.