Abacus.AI ships open-weight Smaug models it says close the gap to frontier agents at a fraction of the cost
Smaug Agentic, Flash and Mini are open-weight models fine-tuned for long-running agent loops, pitched as a 10-100x cheaper substitute for Opus- and GPT-class models in enterprise agentic workflows.
Abacus.AI on September 10 released Smaug, a line of three open-weight models the company fine-tuned specifically for long-running, self-improving agent loops rather than single-turn chat: Smaug Agentic (a 2 trillion-parameter model built on Kimi K3, aimed at complex coding loops), Smaug Flash (built on DeepSeek Flash, aimed at persistent personal agents that connect to WhatsApp, Telegram and Slack), and Smaug Mini (a 27-billion-parameter model for multimodal and smaller reasoning workloads). All three are downloadable on Hugging Face and available through Abacus.AI's RouteLLM API.
The pitch is cost, not raw capability
CEO Bindu Reddy's own framing concedes the gap rather than papering over it -- open-weight models, she said, are "rapidly closing the gap to frontier closed models, but still underperform in long-running agent loops," with Smaug pitched as the fix at a fraction of the price. Abacus.AI says its fine-tuning technique -- combining human-curated agentic traces with synthetic data drawn from difficult examples -- lifts long-running agentic-loop performance 15-20% without added inference cost, and that Smaug Agentic ran 113 DeepSWE coding tasks over more than seven hours with zero infrastructure errors or timeouts, a figure the company published on its own benchmark rather than an independent one. Enterprises can run any of the three models inside their own VPC or on in-house GPU clusters.
This launch is single-sourced to Abacus.AI's own announcement -- no independent outlet with a confirmed byline had reviewed or corroborated the benchmark claims as of publication, so the performance numbers above should be read as the vendor's own reporting, not third-party verification. That gap between a vendor's benchmark and a customer's actual measured outcome is precisely what Merit AC's own scoring is designed to check.