Commit 39ec3fd
feat(gooddata-eval): score dashboard editing alongside creation
The dashboard-skill evaluator scored freshly drafted dashboards only. It now
also scores edits of a saved dashboard, with expected_output.type selecting
everything that differs between the two.
An edit relays two parts: the dashboard as it stood, and an RFC 6902 patch
against it. The patch is applied to a copy of that base and the result is what
every assertion reads. Scoring the base would pass a case whose change never
happened, and skipping the agent's `test` guards would pass a patch written
against a document other than the one it shipped, so the guards are applied
with everything else and a failing one fails the case.
type also decides which skill must have been activated -- dashboard_builder for
creation, dashboard_editor for editing -- and that check now gates rather than
merely being reported. An edit answered by the builder is a routing failure
even when the dashboard that comes back looks plausible.
Three assertions are new to both kinds. saved_dashboard_id must be absent on a
draft and must match the edited dashboard on both the response and the patch.
Widget width is checked wherever the expectation carries it, without which a
resize case asserts nothing. Widget titles gate on an edit, where the
expectation states the title the change must leave behind; on creation they
stay a note, since there the title is the model's own wording.
Attribute filters can now be asserted too, for the case where an edit that was
not asked to touch them drops or narrows them anyway. Both that and the date
range are opt-in: a case carrying neither key is not checked, so the cases that
are about something else are unaffected. The date filter is read from whichever
shape the document uses -- a drafted dashboard hangs one off each tab, a saved
one relayed for editing carries a single map at the root keyed by local
identifier, where the role has to be read from each entry's own type.
A saved dashboard relayed for editing is version 2, carrying sections at the
root and no tabs, so widgets are read from whichever shape the document uses.
The filters score is published as dashboard_filters_correct: alert_skill
already publishes filters_correct, and gdc-nas resolves a trace's skill by
which score names it carries.
jira: QA-29347
risk: low
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>1 parent 9be051b commit 39ec3fd
4 files changed
Lines changed: 936 additions & 165 deletions
File tree
- packages/gooddata-eval
- src/gooddata_eval/core/agentic
- tests
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
13 | 13 | | |
14 | 14 | | |
15 | 15 | | |
| 16 | + | |
16 | 17 | | |
17 | 18 | | |
18 | 19 | | |
| |||
0 commit comments