Challenge
One of our clients serves large enterprise customers whose work requires agents to solve complex problems. Each task had an objectively verifiable outcome, but the sequence of reasoning, tool calls, and actions needed to reach it was not predefined. Different systems could arrive at the same result through radically different trajectories, with significant differences in token usage, latency, reliability, and cost. As demand grew, the client’s AI spending was rising rapidly.
Published token prices could suggest which model was cheapest, while public benchmarks could suggest which was most capable. Team feedback and anecdotes were also valuable for identifying promising systems. None of these signals showed the cost of reliably completing the client’s actual work. Results depended on the complete system: the model, provider, harness, and tools.
The question was not simply whether an agent could complete the work. It was which combination could complete it reliably with unit economics that scaled.
Solution
789 Labs converted the client’s common workloads into a repeatable evaluation suite using an industry-standard framework. These workloads required agents to use purpose-built internal tools that were unique to the client’s business and not natively supported by standard coding agents. Because every task produced a verifiable result, we could measure whether the agent completed the work correctly rather than relying on subjective impressions of its response.
We tested OpenCode, Claude Code, Codex, and a custom harness that treated the client’s tools as a first-class part of the agent loop. These harnesses were evaluated across models and providers, including OpenAI, Anthropic, and OpenRouter. For every configuration, we measured reliability, latency, total cost, token usage, and the trajectory the agent took to complete the work. This revealed the cost of a successful outcome, not merely the advertised price of a token.
The evaluation suite was built to evolve. New models, providers, harness updates, and tools can be added as the market changes, allowing the client to continually identify the most reliable and cost-effective system for its work.
Impact
The benchmarks identified lower-cost systems that could handle the client’s real workloads without sacrificing reliability.
- 75% reduction in API costs while maintaining the required reliability
Are you paying too much for AI?
If your AI spending is growing but you cannot clearly measure the cost of a successful outcome, you may be overpaying. We benchmark your real workloads across models, providers, harnesses, and tools to find the reliability you need with economics that scale. Chat with our team!