Embrasure
Benchmark v0.1July 2026

Enterprise Data Index

A controlled benchmark of frontier models on enterprise data tasks.

Abstract

We evaluate ten frontier models on the same 50-task suite using a fixed Embrasure stack, context snapshot, tools, and inference policy.

The primary endpoint is Verified Task Success Rate. Executed queries must match withheld oracles; unsupported questions are graded on clarification or refusal. The release will include model configurations, task outcomes, 95% intervals, and verification artifacts.

Results under review

Model names and scores are withheld until publication.

The preview leaderboard uses placeholder names and scores. Real results will be published at launch.

Preliminary observation. The largest model was not the highest-scoring model. The best and worst scores differed by 42 percentage points.

Release

The study is currently under review.