Commit fc15e03
feat(gooddata-eval): evaluate the agentic dashboard-creation skill
Adds the agentic_dashboard_skill kind: it drives the conversation until the
agent produces a draft, then scores that draft.
The gate holds only what the agent decided -- the draft call succeeded, the
response carries the expected part, every expected chart is present, every
tab's date filter matches, and it authored at least the required number of
charts. Two observations sit beside the gate and never fail a run: whether a
set_skills call activated dashboard_builder, and whether the response's
references carried every widget id. The second one matters -- gen-ai rejects a
draft naming an unresolvable visualization before it can succeed, so a widget
missing from the references means reference building degraded on the way out,
which is a platform fault and must not be scored as the model's.
A chart the fixture gives an id is matched on that id alone, since the widget
title is title_override or fallback_title and the override is the model's own
choice; a differing title is reported as a note. A chart the fixture marks with
a null id has no identity to match on, so there the title is the match, taken
among the authored charts only -- matching it against every chart would let an
existing one of the same title satisfy it.
The simulated user is deterministic rather than an LLM call: when the agent
asks back instead of drafting, the only things it still needs are the charts
and the date range, and both come straight from the expectation. That reply is
byte-identical every turn, which is why the loop caps at one clarifying round
instead of inheriting metric_skill's seven.
The expectation is validated before the first API call: a fixture with no
visualizations would pass every chart check vacuously, and a date range the
simulated user cannot phrase would otherwise surface only on the branch where
the agent asks back, passing or crashing depending on what the model chose.
jira: QA-29347
risk: low
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>1 parent 1ab844c commit fc15e03
6 files changed
Lines changed: 1278 additions & 0 deletions
File tree
- packages/gooddata-eval
- src/gooddata_eval
- cli
- core/agentic
- tests
Lines changed: 18 additions & 0 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
13 | 13 | | |
14 | 14 | | |
15 | 15 | | |
| 16 | + | |
16 | 17 | | |
17 | 18 | | |
18 | 19 | | |
| |||
40 | 41 | | |
41 | 42 | | |
42 | 43 | | |
| 44 | + | |
43 | 45 | | |
44 | 46 | | |
45 | 47 | | |
| |||
85 | 87 | | |
86 | 88 | | |
87 | 89 | | |
| 90 | + | |
| 91 | + | |
| 92 | + | |
| 93 | + | |
88 | 94 | | |
89 | 95 | | |
90 | 96 | | |
| |||
183 | 189 | | |
184 | 190 | | |
185 | 191 | | |
| 192 | + | |
| 193 | + | |
| 194 | + | |
| 195 | + | |
| 196 | + | |
| 197 | + | |
| 198 | + | |
| 199 | + | |
| 200 | + | |
| 201 | + | |
| 202 | + | |
| 203 | + | |
186 | 204 | | |
187 | 205 | | |
188 | 206 | | |
| |||
Lines changed: 18 additions & 0 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
16 | 16 | | |
17 | 17 | | |
18 | 18 | | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
19 | 29 | | |
20 | 30 | | |
21 | 31 | | |
| |||
62 | 72 | | |
63 | 73 | | |
64 | 74 | | |
| 75 | + | |
65 | 76 | | |
66 | 77 | | |
67 | 78 | | |
| |||
74 | 85 | | |
75 | 86 | | |
76 | 87 | | |
| 88 | + | |
| 89 | + | |
| 90 | + | |
77 | 91 | | |
78 | 92 | | |
79 | 93 | | |
| |||
91 | 105 | | |
92 | 106 | | |
93 | 107 | | |
| 108 | + | |
| 109 | + | |
| 110 | + | |
94 | 111 | | |
95 | 112 | | |
96 | 113 | | |
| |||
99 | 116 | | |
100 | 117 | | |
101 | 118 | | |
| 119 | + | |
102 | 120 | | |
103 | 121 | | |
104 | 122 | | |
| |||
0 commit comments