Benchmarks

An AI Readiness score measures how well a semantic model is prepared for AI. A benchmark measures what AI actually returns from it. BI Pixie writes a set of questions, computes the correct answer to each one with a DAX query against your own data, asks the questions through the AI you choose, and scores the answers. Nothing is hardcoded.

Use a benchmark to establish a baseline before you enable Copilot on a semantic model, to confirm that a round of curation raised accuracy, to compare what an emulated agent achieves against what Copilot or your own data agent returns, and to find which parts of a semantic model produce unreliable answers.

The Five Steps

A benchmark is built on one screen, top to bottom, and each step has a page of its own in this guide:

The benchmark editor showing what gets tested, the AI that answers the questions, the strategy summary, and the saved questions
  1. AI to Test. Choose which AI answers the questions: Power BI Copilot, a Fabric data agent, or a generic AI agent.
  2. Strategy. Set how many questions, which areas of the business, which tables, how demanding, what period, and how many times each question is asked.
  3. Proposed Fields. Review the measures, columns, filter values, and report visuals the questions will be built from.
  4. Questions and Answers. Review, edit, and add questions, and check every correct answer before anything runs.
  5. Run and Results. Run the benchmark, read the result, and repeat it later without changing a thing.

Before You Start

  • Building a benchmark is included on every plan. You can pick the semantic model, choose the AI, generate the questions, and inspect the correct answers without spending anything. A plan buys the run; see Allowances by plan.
  • Fabric access to the semantic model, at Contributor level or higher on its workspace. BI Pixie queries the semantic model with your identity.
  • A Fabric capacity, for two of the three AI options. The query engine behind Copilot in Power BI refuses to answer when the workspace is not on one, and a Fabric data agent needs a paid capacity. Benchmarking through a generic AI agent needs no Fabric capacity. See AI to Test.
  • An AI provider, for natural question phrasing and question variations. Without one, BI Pixie falls back to fixed question templates. The generic AI agent is the one option that requires a provider, because the provider is what answers. See AI Provider.
  • No Fabric install is required for Power BI Copilot or the generic AI agent. Only the Fabric data agent needs the BI Pixie Dashboard, and only an account with data residency switched on needs it for every option.

Where a Benchmark Starts

Every benchmark starts from AI Readiness in the left sidebar of the BI Pixie Portal. New benchmark in the page header asks you to choose the workspace and the semantic model; the benchmark button on a row in Your semantic models, or on a semantic model's own page, starts with the semantic model already chosen. That button's label follows what the semantic model already has, and so does the screen it opens:

The Benchmarks tab listing each semantic model with its benchmark score, trend, and number of runs
You select Because The screen opens on
New benchmark, or Start over You want a different benchmark for this semantic model BI Pixie's default strategy, with every card closed and nothing carried over. If a saved benchmark exists, one line names it and offers Open the saved benchmark. The saved benchmark is replaced only when you generate new questions, and BI Pixie names what will be replaced first.
Resume, or Change the questions You want to change the benchmark you have Your saved questions, open, with the strategy and the field proposal closed above them, each stating its own content.
Run benchmark You want to repeat the benchmark you have One card stating what will run: "BI Pixie asks your semantic model 15 questions through Power BI Copilot and checks every answer against your data", the questions behind a closed disclosure, and Run benchmark as the first control. Change the questions and Start over are offered as the alternatives.

A benchmark you built and did not run is saved automatically, and its row shows a draft line such as "Draft: 14 questions saved 3 days ago" beside a Resume button. Opening any of the three screens and leaving changes nothing you had saved. While a run is underway, the row offers View progress instead.

The portal and the BI Pixie Workload for Microsoft Fabric run the same benchmark, with the same settings and the same results. One detail differs, and it is noted where it applies: the portal's AI provider is usually your own account rather than the BI Pixie data agent.

Allowances by Plan

Every plan can build a benchmark. A plan buys runs, and the screen states your allowance before any questions are generated. Generating never spends a run.

Plan Benchmark runs Notes
Free 1 in total Offered on the account's one tracked semantic model. Before you spend it, the Run step offers See a sample result, a canned result showing what a scored benchmark looks like. It is withdrawn once your own result exists.
Standard 3 in total A lifetime allowance, not a monthly one. The Quick check and Standard question counts are available.
Pro 100 a month Adds the Thorough and Custom question counts.
Enterprise Unlimited Everything in Pro.

A run counts against the allowance identically whichever AI measures it, and deleting a run does not return it. Running a benchmark adds its semantic model to your tracked items, one item per semantic model, stated before the run starts. One benchmark of a given semantic model runs at a time: while one is underway, BI Pixie holds a second start on that semantic model and offers View progress instead. Every strategy setting is available on every plan, and the only tiered one is the question count. Current prices are on the pricing page.

Nothing runs a benchmark for you. A run asks real questions of a live AI, so every benchmark is one you start. Schedules cover AI Readiness assessments only.

Data residency. With data residency switched on, benchmark results are written to your own Fabric lakehouse instead of to BI Pixie's storage. The BI Pixie Dashboard install is therefore required before any benchmark can run, whichever AI you test. The portal states this before you generate the questions. See Data Residency.

What's Next