Skip to content

[bot] Merge master/15ae389b into rel/dev - #1828

Merged
yenkins-admin merged 2 commits into
rel/devfrom
snapshot-master-15ae389b-to-rel/dev
Sep 24, 2026
Merged

yenkins-admin merged 2 commits into
rel/devfrom
snapshot-master-15ae389b-to-rel/dev

Conversation

@yenkins-admin

Copy link
Copy Markdown
Contributor

🚀 Automated PR to perform merge from master into rel/dev with changes up to 15ae389 (created by https://github.com/gooddata/gooddata-python-sdk/actions/runs/35977905980).

FrankHuynh and others added 2 commits September 24, 2026 14:16
Adds the agentic_dashboard_skill kind: it drives the conversation until the
agent produces a draft, then scores that draft.

The gate holds only what the agent decided -- the draft call succeeded, the
response carries the expected part, every expected chart is present, every
tab's date filter matches, and it authored at least the required number of
charts. Three observations sit beside the gate and never fail a run: whether a
set_skills call activated dashboard_builder, whether the response's references
carried every widget id, and whether the widget titles matched. They are scored
rather than only printed, because each of them is a decision to stop failing on
something, and without a score nothing would record how often it happens.

The references one matters most: gen-ai rejects a draft naming an unresolvable
visualization before it can succeed, so a widget missing from the references
means reference building degraded on the way out, which is a platform fault and
must not be scored as the model's.

A chart the fixture gives an id is matched on that id alone, since the widget
title is title_override or fallback_title and the override is the model's own
choice. A chart the fixture marks with a null id has no identity to match on,
so there the title is the match, taken among the authored charts only --
matching it against every chart would let an existing one of the same title
satisfy it.

The simulated user is deterministic rather than an LLM call: when the agent
asks back instead of drafting, the only things it still needs are the charts
and the date range, and both come straight from the expectation. That reply is
byte-identical every turn, so the loop sends it once and stops. Repeating it
answers nothing, and when the agent asked something the expectation does not
cover, repeating it hides the signal that the question was too vague. Two turns
also keep the worst case inside the per-test timeout gdc-nas derives.

The expectation is validated before the first request: a fixture with no
visualizations would pass every chart check vacuously, and a date range the
simulated user cannot phrase would otherwise surface only on the branch where
the agent asks back, passing or crashing depending on what the model chose.

jira: QA-29347
risk: low
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
feat(gooddata-eval): evaluate the agentic dashboard-creation skill
@yenkins-admin
yenkins-admin merged commit 39eadc3 into rel/dev Sep 24, 2026
1 check passed
@yenkins-admin
yenkins-admin deleted the snapshot-master-15ae389b-to-rel/dev branch September 24, 2026 08:54
@coderabbitai

coderabbitai Bot commented Sep 24, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 4dcf9362-73b4-429d-8682-4a3aa4f5d3ef

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Sep 24, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.40893% with 25 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.99%. Comparing base (9927b28) to head (15ae389).
⚠️ Report is 593 commits behind head on rel/dev.

Files with missing lines Patch % Lines
.../src/gooddata_eval/core/agentic/dashboard_skill.py 91.28% 25 Missing ⚠️
Additional details and impacted files
@@             Coverage Diff             @@
##           rel/dev    #1828      +/-   ##
===========================================
+ Coverage    82.88%   82.99%   +0.11%     
===========================================
  Files          328      329       +1     
  Lines        21235    21526     +291     
===========================================
+ Hits         17600    17866     +266     
- Misses        3635     3660      +25     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants