[bot] Merge master/15ae389b into rel/dev - #1828
Conversation
Adds the agentic_dashboard_skill kind: it drives the conversation until the agent produces a draft, then scores that draft. The gate holds only what the agent decided -- the draft call succeeded, the response carries the expected part, every expected chart is present, every tab's date filter matches, and it authored at least the required number of charts. Three observations sit beside the gate and never fail a run: whether a set_skills call activated dashboard_builder, whether the response's references carried every widget id, and whether the widget titles matched. They are scored rather than only printed, because each of them is a decision to stop failing on something, and without a score nothing would record how often it happens. The references one matters most: gen-ai rejects a draft naming an unresolvable visualization before it can succeed, so a widget missing from the references means reference building degraded on the way out, which is a platform fault and must not be scored as the model's. A chart the fixture gives an id is matched on that id alone, since the widget title is title_override or fallback_title and the override is the model's own choice. A chart the fixture marks with a null id has no identity to match on, so there the title is the match, taken among the authored charts only -- matching it against every chart would let an existing one of the same title satisfy it. The simulated user is deterministic rather than an LLM call: when the agent asks back instead of drafting, the only things it still needs are the charts and the date range, and both come straight from the expectation. That reply is byte-identical every turn, so the loop sends it once and stops. Repeating it answers nothing, and when the agent asked something the expectation does not cover, repeating it hides the signal that the question was too vague. Two turns also keep the worst case inside the per-test timeout gdc-nas derives. The expectation is validated before the first request: a fixture with no visualizations would pass every chart check vacuously, and a date range the simulated user cannot phrase would otherwise surface only on the branch where the agent asks back, passing or crashing depending on what the model chose. jira: QA-29347 risk: low Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
feat(gooddata-eval): evaluate the agentic dashboard-creation skill
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## rel/dev #1828 +/- ##
===========================================
+ Coverage 82.88% 82.99% +0.11%
===========================================
Files 328 329 +1
Lines 21235 21526 +291
===========================================
+ Hits 17600 17866 +266
- Misses 3635 3660 +25 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
🚀 Automated PR to perform merge from master into rel/dev with changes up to 15ae389 (created by https://github.com/gooddata/gooddata-python-sdk/actions/runs/35977905980).