Strategy
The Benchmark strategy card decides which questions are written, how they are asked, and how later runs are graded. It works the same way whichever AI you test. The card opens closed, with a one-line summary of the current settings, such as "15 questions · Last month · All domains · Usage-first tables · Includes undocumented fields · Balanced mix · Asked 3 different ways". Select Define strategy to open the dials, or Edit strategy where a saved benchmark already exists, and Collapse to fold them away. The strategy is saved with the benchmark, one per semantic model.
Every dial has a default that suits most semantic models, so the shortest path is to open the dials, select Propose fields, and generate. Every dial is available on every plan. The only tiered setting is the question count, where the Thorough and Custom options wear a Pro pill until your plan includes them.
The Dials
| Setting | Options | What it changes |
|---|---|---|
| How many questions | Quick check (about 10), Standard (about 15, the default), Thorough (about 30), or Custom (3 to 100) | How many questions the benchmark asks. More questions give a more reliable score and a longer run, and each question is real work for the AI that answers it. Thorough and Custom are included with the Pro plan and above. |
| Reporting period | Use the last complete reporting period (the default, with a Period size of Last month or Last year), Set the period myself, or Keep dates dynamic | The date range the questions cover. A fixed period means new data is much less likely to move your score between runs. Dynamic dates mean the questions always ask about the data as it stands on the day you run the benchmark, so scores from different runs are not directly comparable. The card names the period it resolved to, and warns before you generate if that period falls outside the range your data covers, or if some tables hold no data in it. |
| Business domains | All domains, a selection of the detected domains, or a domain you add yourself with the tables it covers | Which areas of the business the questions come from. BI Pixie reads the domains from the semantic model's AI Instructions through your AI provider, and writes questions only about the areas you select, using the tables and fields those areas cover. A semantic model with no AI Instructions has no domains to offer, and the card says so. |
| Tables to focus on | Top tables by real usage (recommended), with a Usage focus percentage, Choose tables myself, or Spread across the whole semantic model | Which tables and fields the questions are built from. Under the recommended option, the usage focus percentage sets how many questions target your most-used tables and fields, with the rest exploring the wider semantic model. Usage focuses the questions and never limits them: where your reports are not tracked, BI Pixie ranks by structure instead. |
| Question complexity | Simple check (70% simple, 30% standard), Balanced (40% simple, 40% standard, 20% complex, the default), Stress test (20% simple, 40% standard, 40% complex), or Custom | How demanding the questions are. A simple question uses one measure. A standard question groups one measure by one column. A complex question combines two measures with a filter your users apply, or asks for a top value. Every question in the review list carries its level. |
| Include undocumented fields | On (the default) or off | Whether fields with no description are eligible for questions. Documented fields rank first either way. Leaving it on shows where descriptions would help; turning it off restricts the benchmark to documented fields. |
| How many times to ask each question | 1 to 5, with 3 as the default | Asking the same question more than once is how BI Pixie finds answers that change from one attempt to the next, reported as Unstable. The card states the total number of asks the run will send. |
| How to word each question when it is asked again | Ask each question a different way (the default) or Ask each question the same way | Available from two asks upward, and requires an AI provider. Asking a different way has BI Pixie write question variations that ask for the same answer in other words, so a difference in the answers means accuracy depends on phrasing. Asking the same way sends identical text, so a difference means the AI is not answering consistently. You can review every variation before the run. |
| Which correct answers to use when you run this benchmark again | The original answers (the default) or Answers recalculated from your data | BI Pixie saves the correct answers when the benchmark is created. Scoring later runs against those saved answers keeps runs comparable, and the result reports any question whose underlying data has since changed. Scoring against current data re-runs each question's DAX query before the run, so the answers always match your data on the day. The setting applies to later runs, and past runs are never rescored. The same control appears on every result. |
Running Without an AI Provider
Without a provider, the benchmark still builds and runs, using fixed question templates. Business-domain detection and question variations are unavailable, and each states that where it appears. The generic AI agent cannot run at all, because the provider is what answers. See AI Provider.
What's Next
- Proposed Fields, what Propose fields shows you.
- AI to Test, the card above this one.