Inboxsmith

AI news videos

NVIDIA NeMo Switchyard: LangChain Reports 74% Lower AI Cost

NVIDIA announced NeMo Switchyard on August 11, 2026, an open source model routing library for AI agents, and published what its early enterprise partners measured with it. Everything below comes from NVIDIA's own blog post by Kari Briski.

Watch on YouTube

Transcript

NVIDIA announced today that Ramp, LangChain and Cognition cut agent costs sharply by routing work between models.

LangChain reported seventy four percent lower cost across one hundred forty five multi turn tasks, sending just seven percent of calls to a frontier model, accepting six percent lower accuracy.

Ramp said Switchyard matched a frontier model while cutting cost fifty eight percent and runtime thirty three percent. Cognition put it inside Devin Desktop, mean cost down twenty eight percent.

Boomi reported one hundred percent domain routing accuracy. CrowdStrike, Harvey and CodeRabbit are customizing the model. Every figure comes from NVIDIA or its partners, with no independent verification.

Inboxsmith helps small businesses handle calls and messages so nothing gets missed. Please like and subscribe for more news.

Sources

Every claim in this video comes from the top ranking coverage of this topic. The claims and where each one came from:

  • LangChain: With NeMo Switchyard, achieved 74% lower cost in 145 multi-turn Deep Agents tasks by routing only 7% of calls to a frontier model, at a 6% accuracy tradeoff.(NVIDIA's official announcement)
  • Ramp: Used NeMo Switchyard to match a frontier model's performance while cutting costs by 58% and runtime by 33% in Ramp SWE-Bench.(NVIDIA's official announcement)
  • Cognition: Integrated the NVIDIA NeMo Switchyard staged router into Devin Desktop for NVIDIA internal use, achieving near-frontier performance on FrontierCode Main while reducing mean cost by 28% relative to routing all requests to a single underlying frontier model.(NVIDIA's official announcement)
  • Boomi: Evaluated Switchyard across five routing capabilities, achieving 100% domain-routing accuracy, sending 59% of traffic to a 5x faster fine-tuned model and reducing later-turn latency by 21%.(NVIDIA's official announcement)
  • AI leaders across industries are customizing Nemotron 3.5 Lightning for their workloads, including CrowdStrike for cybersecurity, Harvey with Trajectory for legal services and CodeRabbit with Baseten for code review, helping improve accuracy for domain-specific agentic tasks.(NVIDIA's official announcement)
  • NVIDIA internal benchmarks show that NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone.(NVIDIA's official announcement)
  • Also, NVIDIA is releasing NeMo Switchyard, an open source library for smart routing inside popular agent tools.(NVIDIA's official announcement)
  • Cadence: Improved efficiency by 9.9% by using the ChipStack AI Super Agent for a formal verification use case.(NVIDIA's official announcement)

We make Inboxsmith.

An AI receptionist that never misses a business call.

See how it works