Enterprise Data Index
We evaluated frontier models on 50 enterprise data questions spanning CRM, billing, product usage, support, and identity. Answers were graded by executed results, not SQL similarity or LLM judgment.
GPT-5.6 Sol scored highest.
Its observed Verified Task Success Rate was 58%. Kimi K3 scored 38% in its additional run.
GPT-5.6 Sol
29 of 50 tasks$23.98 provider cost
58%
95 percent interval 46 to 68 percent.Claude Opus 4.8
25 of 50 tasks$18.48 provider cost
50%
95 percent interval 38 to 62 percent.Kimi K3
19 of 50 tasks$8.21 provider costAdditional run
38%
95 percent interval 26 to 50 percent.GLM-5.2
18 of 50 tasks$0.37 provider cost
36%
95 percent interval 24 to 48 percent.Gemini 3.1 Pro Preview
8 of 50 tasks$5.43 provider cost
16%
95 percent interval 8 to 26 percent.
What this run measures
Tasks require models to select the right source, apply business definitions, reason across systems and time, clarify ambiguous requests, and abstain when the available data cannot support an answer. The benchmark runs against a deterministic synthetic B2B SaaS company.
How the run worked
- Questions
- 50 fixed tasks over a deterministic synthetic B2B SaaS company
- Task mix
- 45 query answers, 3 clarifications, and 2 supported refusals
- Run
- One isolated attempt per model and task; Kimi K3 was run separately
- Grading
- Executed results compared with deterministic reference results
- Context
- The same AgentLake configuration, tools, and execution limits for every model
These results describe models running within a fixed AgentLake system. They do not compare Databricks Genie, Snowflake CoCo, Claude Code, or Codex as complete products. The methodology, task set, and normalized results will be released after the remaining review and conformance checks are complete.
Why we built it
Reliable enterprise data work depends on model quality, business context, execution policy, and evidence. AgentLake holds those conditions fixed so an answer can be checked rather than taken on confidence.
