Skip to main content

Eval Metrics

Metrics score each (golden, arm, repeat) run. They are declared in a suite's metrics: list and resolved by name — no Python needed. Each metric produces a score from 0.0 to 1.0 per criterion; some metrics are only applicable when the golden carries the matching reference data, in which case they are skipped (and skipped LLM calls cost nothing).

The overall score for a run is the mean of all applicable criteria. The flat metric-store row (and the scorecard) treat a run as passing when the overall score meets the suite's threshold — the threshold of the suite's first metric.

metrics:
- type: schema_membership
threshold: 1.0
- type: reference_value_match
- type: llm_judge
name: sql_soundness
threshold: 0.5
params:
evaluation_params: [input, predicted_sqls]
criteria: "Is the SQL structurally sound and a plausible way to answer the question?"

Common Parameters​

ParameterDefaultDescription
type—Metric type (see below)
namethe typeDisplay name; required to use the same type (e.g. llm_judge) more than once in a suite
threshold0.5Score threshold for the criterion
params—Type-specific configuration (used by llm_judge)

Deterministic Metrics​

No LLM calls, zero cost, fully reproducible.

TypeWhat it checksApplicable when
schema_membershipEvery table and (best-effort) table.column referenced in the agent's SQL exists in the semantic layer's schema catalogThe trace contains SQL
expected_view_recallRecall over golden.expected_views — did the agent touch the views it needed to? Defensive extra fetches are not penalizedexpected_views is set
reference_value_matchThe agent's final query results match golden.reference_values within tolerance — catches "queried the adjacent view, got a plausible but wrong number"reference_values is set
step_efficiencyDuplicate tool-call (loop) detection: 1 - duplicates / total tool callsAlways
hit_limitPass/fail on whether the agent hit its turn budget without submitting an answer — keeps "agent gave up" a first-class, visible categoryAlways
metrics:
- type: schema_membership
threshold: 1.0
- type: expected_view_recall
- type: reference_value_match
- type: step_efficiency
threshold: 0.8
- type: hit_limit

LLM Judge​

llm_judge is a generic declarative LLM-as-judge metric: the rubric lives entirely in params, so multiple judge criteria (answer correctness, SQL soundness, groundedness, ...) can be defined in one suite without any new code. The judge is asked directly for a 0.0–1.0 score via structured output.

params keyRequiredDescription
criteriaone ofA single rubric statement. Mutually exclusive with evaluation_steps
evaluation_stepsone ofAn ordered list of rubric steps the judge must follow in order
evaluation_paramsNo (default [input, final_answer])Which context fields to render for the judge — see below
judge_modelNo (default anthropic:claude-sonnet-5)Model used for judging
applicable_whenNoA context field name; the metric (and its LLM call) is skipped unless that field is truthy on the golden/trace

Available evaluation_params values:

ValueSource
inputThe golden's question
final_answerThe agent's final natural-language answer
predicted_sqlsAll SQL the agent ran
query_resultsRows returned by the final query
reference_answer_textThe golden's reference answer
tool_callsThe agent's tool-call history
schema_catalogThe semantic layer's view names
widgetsChart/visualization metadata the agent produced (AgentTrace.widgets — chart type, columns, row count, title)
is_dashboard_turnWhether this trace came from a dashboard-building turn (AgentTrace.is_dashboard_turn)

Judge-prompt versioning is built in: every score is stamped with a hash of the rendered judge prompt, so a rubric edit is never silently compared against results scored under the old rubric. Use weiser eval-calibrate to verify a judge agrees with human labels before trusting it.

metrics:
- type: llm_judge
name: sql_soundness
threshold: 0.7
params:
criteria: "Is the SQL structurally sound given the schema, even without live execution?"
evaluation_params: [predicted_sqls, schema_catalog]
judge_model: anthropic:claude-sonnet-5

Ready-Made Presets​

Three documented, reusable shapes for the common judge criteria (the mechanical parts — schema membership, expected-view recall, reference values — are already covered by the deterministic metrics, so these stay scoped to what is genuinely subjective):

Answer correctness — compare the agent's substantive claims against a reference answer. Applicable only when the golden has reference_answer_text:

- type: llm_judge
name: answer_correctness
threshold: 0.7
params:
applicable_when: reference_answer_text
evaluation_params: [input, final_answer, reference_answer_text]
criteria: >
Compare the agent's substantive claims (numbers, trends, named entities)
against the reference answer. Score 0.0 for a material factual or numeric
contradiction, 1.0 when the substance matches even if phrasing differs. Do
not penalize stylistic differences, extra correct detail, or a different but
equally valid framing of the same facts.

SQL soundness — is the submitted SQL structurally sound and plausible, given the schema?

- type: llm_judge
name: sql_soundness
threshold: 0.5
params:
evaluation_params: [input, predicted_sqls, schema_catalog]
criteria: >
Given the question and the known schema, is the submitted SQL structurally
sound and a plausible way to answer the question? Judge plausibility only --
you do not have live execution results here. Do not penalize style choices
(aliasing, formatting, CTE vs. subquery).

Groundedness — does the final answer accurately reflect the query results, without fabricating or omitting material facts?

- type: llm_judge
name: groundedness
threshold: 0.7
params:
evaluation_params: [final_answer, query_results]
criteria: >
Does the final natural-language answer accurately reflect the query_results
returned, without fabricating or omitting material facts? This checks
internal consistency only -- it does not check whether query_results are
themselves correct.

Chart type appropriateness — is the chosen chart type reasonable for the query result shape? Applicable only when the trace has at least one widget (AgentTrace.widgets). Optionally accepts a chart_rules argument when built via chart_type_appropriateness_metric_config(chart_rules=...), so the judge can be handed the app's own chart-agent's exact selection rules instead of generic taste:

- type: llm_judge
name: chart_type_appropriateness
threshold: 0.6
params:
applicable_when: widgets
evaluation_params: [input, widgets]
criteria: >
Given the query result shape (columns, row count) and the user's question, is
the chosen chart type reasonable? Do not dock points for "a different acceptable
option would have been more insightful" when several chart types are all
defensible for the same data shape.

Dashboard composition — from the tool calls that mutated a dashboard, is the resulting structure sensible? Applicable only on dashboard-building turns (AgentTrace.is_dashboard_turn):

- type: llm_judge
name: dashboard_composition
threshold: 0.6
params:
applicable_when: is_dashboard_turn
evaluation_params: [input, tool_calls]
criteria: >
From the tool calls that mutated the dashboard, is the resulting structure
sensible -- no obviously redundant charts, reasonable organization given what
the user asked for?

LLM-judge metrics cost a real LLM call per question — comment them out for a fully deterministic, zero-cost run.

Combining Metrics​

A typical suite mixes cheap deterministic gates with one or two subjective judge criteria:

metrics:
- type: schema_membership
threshold: 1.0
- type: expected_view_recall
- type: reference_value_match
- type: step_efficiency
threshold: 0.5
- type: hit_limit
- type: llm_judge
name: sql_soundness
threshold: 0.5
params:
evaluation_params: [input, predicted_sqls]
criteria: >
Given the question, is the submitted SQL structurally sound and a
plausible way to answer it? Do not penalize style choices.