Run and Results
Run the Benchmark
Select Run benchmark. The button states how many questions will be asked, and the first run on a semantic model states that it adds the semantic model to your plan's tracked items. Every run executes in BI Pixie's service, with live progress through three phases: working out the correct answers, asking the questions, and scoring the answers. Where the answers come from depends on the AI you test: a Fabric data agent answers inside your Fabric tenant on your capacity, Power BI Copilot runs its queries on the capacity behind the semantic model, and a generic AI agent answers through your AI provider with no Fabric capacity involved. A run usually takes a few minutes.
You do not have to watch. The run continues after you leave the page, the result is saved to the semantic model's history, and the row in Your semantic models offers View progress until it finishes. If a run takes unusually long, a warning offers Stop watching, which detaches the page without canceling the run.
On the free plan, the Run step offers See a sample result before you spend your single run: a canned result that shows what a scored benchmark looks like, with no measurement of your semantic model in it. Once your own result exists, the sample is withdrawn, so a scored result on screen is always yours.
Read the Result
The result opens under the breadcrumb AI Readiness / the semantic model / Benchmark and the date, headed "Sales Analytics, checked Aug 20, 2026 · via Power BI Copilot". Every score names what measured it.
The score is correct answers out of every question asked. A benchmark that asked 12 questions and answered 7 correctly reads "7 of 12 questions answered correctly", whether or not the other 5 reached a verdict. A question that could not be scored costs exactly what a wrong one costs, so a run that lost questions to a busy capacity can never look better than a run that answered them all. A run in which no question reached a verdict reads Not scored rather than 0%.
About this run
Under the score, a labeled row for each fact about the run, each opening to a plain explanation of what it means for the score:
- Answered by. The AI that answered, with its provider and, where known, its AI model.
- Questions. How many were asked and how many were scored.
- What it covers and Question mix. The business domains and the simple, standard, and complex split.
- Reporting period. The date range the questions covered, so a past score can be read in context.
- What the AI could see. On the generic AI agent, exactly what was assembled for it, including how many AI-excluded fields were withheld and whether the description was shortened.
- Correct answers. Whether the run was scored against the answers saved with the benchmark or against answers recalculated on the day.
Verdicts
Each question receives one verdict: Correct, Unstable (correct on some asks and wrong on others, which asking more than once exists to expose), Wrong, or Not scored, with the reason. The reason names the right culprit: the AI declining to answer, a busy capacity, a ground truth that returned nothing, or BI Pixie itself being unable to read the reply.
Expand a question for the full detail: The agent answered, The right answer from your data, and The queries behind this result, holding the query the AI wrote beside the ground-truth DAX BI Pixie used to check it. A Fabric data agent writes and runs its own query inside Fabric and hands back only its answer, so on those runs the section shows BI Pixie's query alone and a note says why.
Answer checking
Every answer is first compared directly against the correct value computed from your data. Where an AI provider is available, an AI check supports that comparison, whichever AI you test: where the wording makes an answer unclear, the check marks the question not scored rather than wrong. It never overrides an answer already scored correct.
Weak spots
A Weak spots section gathers the questions answered wrongly or unstably, each stating where it came from, so you can see where the semantic model falls short. The next step is to improve the semantic model through its Optimizations, then run the benchmark again. Edit questions on the result opens the saved benchmark for changes; editing starts nothing and uses no run.
Repeat the Benchmark
A repeat answers "did my fixes move the score?", which is only meaningful when the second run asked the same questions of the same semantic model through the same AI the same number of times. Run benchmark on the semantic model's row opens a screen stating exactly that, with Run benchmark as its first control; see Where a Benchmark Starts. An opened result offers the same repeat as Run this benchmark again. The result of a repeat states the comparison in words, such as "Up from 58% on Aug 1".
- Correct answers. A repeat is graded according to the Which correct answers to use setting, which the result card also carries as The original answers or Answers recalculated from your data. Where the data behind a question has changed since the answers were saved, the result says so and offers Update answers to current data.
- Different AIs are never one trend. Runs are grouped under the AI that produced them, and changing the AI on a semantic model with history asks you to confirm. Runs recorded before this choice existed are labeled as the Fabric data agent.
- A run that cannot be repeated exactly, such as one recorded before BI Pixie stored benchmark questions, says so and offers a new benchmark instead of a repeat that would ask different questions.
- Aged-out periods. When a fixed reporting period has fallen behind your data, the result offers to move to the current last period, which builds a new benchmark and therefore starts a new baseline, or to set the period yourself.
History
Every run is stored under its semantic model. The Benchmarks tab of Your semantic models lists the latest score, the AI it was measured through, its change, and the run count for every semantic model that has runs, and expands to the full list. The semantic model's own page holds the same list in its Benchmarks section, and the New benchmark screen lists Recent benchmark runs while no run is in flight. Each record shows its score on the same bands as the registry, the verdict breakdown behind it, and its change against the run before. Opening a stored run starts nothing and requires no Fabric permissions.
Delete this run removes a single run after confirmation. Delete all results, in the actions menu on the semantic model's page, removes its assessments and benchmark runs together. Deleting a run does not return it to your plan's allowance, and deleting a semantic model's results does not reduce your tracked items. With data residency switched on, results live in your own lakehouse, and BI Pixie deletes from it only with your own storage permission.
What's Next
- Optimizations, to raise the score before you run again.
- Benchmarks, for allowances and the three ways of arriving.
- Set up BI Pixie Dashboard, needed to test a Fabric data agent.
- Data Residency, if results should be written to your own lakehouse.