Model ML Says GPT-5.6 Sol Cuts an Hour of Finance Work to Five Minutes
OpenAI published a customer story on August 10, 2026 about Model ML, a startup whose agents carry finance work from a brief through research and analysis to a finished, editable PowerPoint deck or Excel workbook with traceable sources. Everything below comes from that page, and the benchmark numbers in it are Model ML's own internal evaluation, not an independent test.
Transcript
A finance task that took an hour now takes about five minutes. Model ML credits GPT five point six Sol.
Model ML builds agents for finance teams. From a brief, they run research and analysis, then produce a finished, editable PowerPoint deck or Excel workbook with traceable sources.
In Model ML's own tests, Sol was ready for substantive review far more often than Opus five, leading by sixteen point six percentage points, and produced a deck every time.
Sol used twenty one percent fewer tokens per deck than Fable five, and thirty six percent fewer per workbook than Opus five. Still, the results were mixed on Excel accuracy.
Inboxsmith helps small businesses handle calls and messages so nothing gets missed. Please like and subscribe for more news.
Sources
Every claim in this video comes from the top ranking coverage of this topic. The claims and where each one came from:
- At one global asset manager, a bespoke tearsheet that took an analyst about an hour to assemble now takes about five minutes.(OpenAI's official announcement)
- Model ML's agents help finance professionals carry a workflow from the initial request through research, analysis, and a finished deck or workbook.(OpenAI's official announcement)
- Model ML's own document tooling creates native, editable PowerPoint and Excel files with traceable sources.(OpenAI's official announcement)
- In Model ML's Composite eval, GPT-5.6 Sol cleared the professional-readiness gate, a measure of whether output was ready for substantive review, in 43.3% of cases versus 26.7% for Opus 5, a 16.6 percentage point lead.(OpenAI's official announcement)
- GPT-5.6 Sol completed the PowerPoint workflow and yielded a .pptx file in 100% of test cases, compared with 76% for Opus 5.(OpenAI's official announcement)
- For PowerPoint workflows, GPT-5.6 Sol used about 21% fewer tokens per deck than Fable 5.(OpenAI's official announcement)
- In an Excel workflow, GPT-5.6 Sol used 36% fewer tokens per workbook than Opus 5, 2.44M versus 3.83M.(OpenAI's official announcement)
- Results were mixed on Excel accuracy: GPT-5.6 Sol got every key output right in 50.0% of items versus 60.0% for Opus 5, and aggregate visual quality was 77.9% versus 79.3% for Opus 5.(OpenAI's official announcement)
- The benchmark is Model ML's Composite, the company's own evaluation benchmark for AI in financial services, published by OpenAI on August 10, 2026.(OpenAI's official announcement)
