Enterprise Data Index
A controlled benchmark of frontier models on enterprise data tasks.
Abstract
We evaluate ten frontier models on the same 50-task suite using a fixed Embrasure stack, context snapshot, tools, and inference policy.
The primary endpoint is Verified Task Success Rate. Executed queries must match withheld oracles; unsupported questions are graded on clarification or refusal. The release will include model configurations, task outcomes, 95% intervals, and verification artifacts.
Results under review
Model names and scores are withheld until publication.
The preview leaderboard uses placeholder names and scores. Real results will be published at launch.
Preliminary observation. The largest model was not the highest-scoring model. The best and worst scores differed by 42 percentage points.
Release
The study is currently under review.