Embrasure

Enterprise Data Index

We evaluated frontier models on 50 enterprise data questions spanning CRM, billing, product usage, support, and identity. Answers were graded by executed results, not SQL similarity or LLM judgment.

GPT-5.6 Sol scored highest.

Its observed Verified Task Success Rate was 58%. Kimi K3 scored 38% in its additional run.

  1. GPT-5.6 Sol

    29 of 50 tasks$23.98 provider cost

    58%

    95 percent interval 46 to 68 percent.
  2. Claude Opus 4.8

    25 of 50 tasks$18.48 provider cost

    50%

    95 percent interval 38 to 62 percent.
  3. Kimi K3

    19 of 50 tasks$8.21 provider costAdditional run

    38%

    95 percent interval 26 to 50 percent.
  4. GLM-5.2

    18 of 50 tasks$0.37 provider cost

    36%

    95 percent interval 24 to 48 percent.
  5. Gemini 3.1 Pro Preview

    8 of 50 tasks$5.43 provider cost

    16%

    95 percent interval 8 to 26 percent.
Bars show Verified Task Success Rate. Thin lines show the descriptive 95% task-bootstrap interval from this task set. This single run does not establish a stable ranking. Costs include model and diagnostic judge calls for each 50-question run, rounded to cents.

What this run measures

Tasks require models to select the right source, apply business definitions, reason across systems and time, clarify ambiguous requests, and abstain when the available data cannot support an answer. The benchmark runs against a deterministic synthetic B2B SaaS company.

How the run worked

Questions
50 fixed tasks over a deterministic synthetic B2B SaaS company
Task mix
45 query answers, 3 clarifications, and 2 supported refusals
Run
One isolated attempt per model and task; Kimi K3 was run separately
Grading
Executed results compared with deterministic reference results
Context
The same AgentLake configuration, tools, and execution limits for every model

These results describe models running within a fixed AgentLake system. They do not compare Databricks Genie, Snowflake CoCo, Claude Code, or Codex as complete products. The methodology, task set, and normalized results will be released after the remaining review and conformance checks are complete.

Why we built it

Reliable enterprise data work depends on model quality, business context, execution policy, and evidence. AgentLake holds those conditions fixed so an answer can be checked rather than taken on confidence.