Skip to content

Commit 39ec3fd

Browse files
FrankHuynhclaude
andcommitted
feat(gooddata-eval): score dashboard editing alongside creation
The dashboard-skill evaluator scored freshly drafted dashboards only. It now also scores edits of a saved dashboard, with expected_output.type selecting everything that differs between the two. An edit relays two parts: the dashboard as it stood, and an RFC 6902 patch against it. The patch is applied to a copy of that base and the result is what every assertion reads. Scoring the base would pass a case whose change never happened, and skipping the agent's `test` guards would pass a patch written against a document other than the one it shipped, so the guards are applied with everything else and a failing one fails the case. type also decides which skill must have been activated -- dashboard_builder for creation, dashboard_editor for editing -- and that check now gates rather than merely being reported. An edit answered by the builder is a routing failure even when the dashboard that comes back looks plausible. Three assertions are new to both kinds. saved_dashboard_id must be absent on a draft and must match the edited dashboard on both the response and the patch. Widget width is checked wherever the expectation carries it, without which a resize case asserts nothing. Widget titles gate on an edit, where the expectation states the title the change must leave behind; on creation they stay a note, since there the title is the model's own wording. Attribute filters can now be asserted too, for the case where an edit that was not asked to touch them drops or narrows them anyway. Both that and the date range are opt-in: a case carrying neither key is not checked, so the cases that are about something else are unaffected. The date filter is read from whichever shape the document uses -- a drafted dashboard hangs one off each tab, a saved one relayed for editing carries a single map at the root keyed by local identifier, where the role has to be read from each entry's own type. A saved dashboard relayed for editing is version 2, carrying sections at the root and no tabs, so widgets are read from whichever shape the document uses. The filters score is published as dashboard_filters_correct: alert_skill already publishes filters_correct, and gdc-nas resolves a trace's skill by which score names it carries. jira: QA-29347 risk: low Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1 parent 9be051b commit 39ec3fd

4 files changed

Lines changed: 936 additions & 165 deletions

File tree

‎packages/gooddata-eval/pyproject.toml‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -13,6 +13,7 @@ requires-python = ">=3.10"
1313
dependencies = [
1414
"gooddata-sdk~=1.75.0",
1515
"httpx>=0.27,<1.0",
16+
"jsonpatch>=1.33,<2.0",
1617
"orjson>=3.9.15,<4.0.0",
1718
"pydantic>=2.6,<3.0",
1819
"rich>=13.0,<15.0",

0 commit comments

Comments
 (0)