Benchmarks
An AI Readiness score measures how well a semantic model is prepared for AI. A benchmark measures what AI actually returns from it. BI Pixie writes a set of questions, computes the correct answer to each one with a DAX query against your own data, asks the questions through the AI you choose, and scores the answers. Nothing is hardcoded.
Use a benchmark to establish a baseline before you enable Copilot on a semantic model, to confirm that a round of curation raised accuracy, to compare what an emulated agent achieves against what Copilot or your own data agent returns, and to find which parts of a semantic model produce unreliable answers.
The Five Steps
A benchmark is built on one screen, top to bottom, and each step has a page of its own in this guide:
- AI to Test. Choose which AI answers the questions: Power BI Copilot, a Fabric data agent, or a generic AI agent.
- Strategy. Set how many questions, which areas of the business, which tables, how demanding, what period, and how many times each question is asked.
- Proposed Fields. Review the measures, columns, filter values, and report visuals the questions will be built from.
- Questions and Answers. Review, edit, and add questions, and check every correct answer before anything runs.
- Run and Results. Run the benchmark, read the result, and repeat it later without changing a thing.
Before You Start
- Building a benchmark is included on every plan. You can pick the semantic model, choose the AI, generate the questions, and inspect the correct answers without spending anything. A plan buys the run; see Allowances by plan.
- Fabric access to the semantic model, at Contributor level or higher on its workspace. BI Pixie queries the semantic model with your identity.
- A Fabric capacity, for two of the three AI options. The query engine behind Copilot in Power BI refuses to answer when the workspace is not on one, and a Fabric data agent needs a paid capacity. Benchmarking through a generic AI agent needs no Fabric capacity. See AI to Test.
- An AI provider, for natural question phrasing and question variations. Without one, BI Pixie falls back to fixed question templates. The generic AI agent is the one option that requires a provider, because the provider is what answers. See AI Provider.
- No Fabric install is required for Power BI Copilot or the generic AI agent. Only the Fabric data agent needs the BI Pixie Dashboard, and only an account with data residency switched on needs it for every option.
Where a Benchmark Starts
Every benchmark starts from AI Readiness in the left sidebar of the BI Pixie Portal. New benchmark in the page header asks you to choose the workspace and the semantic model; the benchmark button on a row in Your semantic models, or on a semantic model's own page, starts with the semantic model already chosen. That button's label follows what the semantic model already has, and so does the screen it opens:
| You select | Because | The screen opens on |
|---|---|---|
| New benchmark, or Start over | You want a different benchmark for this semantic model | BI Pixie's default strategy, with every card closed and nothing carried over. If a saved benchmark exists, one line names it and offers Open the saved benchmark. The saved benchmark is replaced only when you generate new questions, and BI Pixie names what will be replaced first. |
| Resume, or Change the questions | You want to change the benchmark you have | Your saved questions, open, with the strategy and the field proposal closed above them, each stating its own content. |
| Run benchmark | You want to repeat the benchmark you have | One card stating what will run: "BI Pixie asks your semantic model 15 questions through Power BI Copilot and checks every answer against your data", the questions behind a closed disclosure, and Run benchmark as the first control. Change the questions and Start over are offered as the alternatives. |
A benchmark you built and did not run is saved automatically, and its row shows a draft line such as "Draft: 14 questions saved 3 days ago" beside a Resume button. Opening any of the three screens and leaving changes nothing you had saved. While a run is underway, the row offers View progress instead.
Allowances by Plan
Every plan can build a benchmark. A plan buys runs, and the screen states your allowance before any questions are generated. Generating never spends a run.
| Plan | Benchmark runs | Notes |
|---|---|---|
| Free | 1 in total | Offered on the account's one tracked semantic model. Before you spend it, the Run step offers See a sample result, a canned result showing what a scored benchmark looks like. It is withdrawn once your own result exists. |
| Standard | 3 in total | A lifetime allowance, not a monthly one. The Quick check and Standard question counts are available. |
| Pro | 100 a month | Adds the Thorough and Custom question counts. |
| Enterprise | Unlimited | Everything in Pro. |
A run counts against the allowance identically whichever AI measures it, and deleting a run does not return it. Running a benchmark adds its semantic model to your tracked items, one item per semantic model, stated before the run starts. One benchmark of a given semantic model runs at a time: while one is underway, BI Pixie holds a second start on that semantic model and offers View progress instead. Every strategy setting is available on every plan, and the only tiered one is the question count. Current prices are on the pricing page.
Nothing runs a benchmark for you. A run asks real questions of a live AI, so every benchmark is one you start. Schedules cover AI Readiness assessments only.
What's Next
- AI to Test, the first choice on the screen.
- Optimizations, to raise the score between runs.
- Set up BI Pixie Dashboard, needed to test a Fabric data agent.
- Add Pixies to Your Reports, so questions are ranked by real usage.