Skip to content

Commit 585abcb

Browse files
committed
feat(gooddata-eval): add data obfuscation evaluator
Add the agentic_obfuscation kind, which checks that canaries are masked in the stored conversation and the Langfuse trace. jira: QA-29442 risk: low
1 parent 9be051b commit 585abcb

22 files changed

Lines changed: 2257 additions & 18 deletions

‎packages/gooddata-eval/README.md‎

Lines changed: 54 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -136,7 +136,7 @@ gd-eval run \
136136
| `--reasoning-effort LEVEL` | server default | `LOW`, `MEDIUM` or `HIGH`, sent as `options.reasoningEffort` on every chat message. Requires the `enableGenAiReasoningEffort` feature flag on the target organization — without it the server ignores the value. Applies to chat items only; `dashboard_summary` items go through the summary endpoint, which has no such option. |
137137

138138
**Concurrency and workspace safety.** Agentic kinds that create workspace objects
139-
(`agentic_metric_skill`, `agentic_alert_skill`, `agentic_conversation`, `agentic_kda_skill`) always run one at a
139+
(`agentic_metric_skill`, `agentic_alert_skill`, `agentic_conversation`, `agentic_kda_skill`, `agentic_obfuscation`) always run one at a
140140
time whatever `--concurrency` says — a metric or alert created and dropped mid-run would otherwise be visible to
141141
another item reading the same catalog. **That protection is for the agentic kinds only:** the single-turn
142142
`metric_skill` and `alert_skill` kinds are still fanned out and the agent performs the same server-side writes on
@@ -535,7 +535,8 @@ A dataset is a folder of `.json` files, one per question:
535535
```
536536

537537
Supported `test_kind` values: `visualization`, `metric_skill`, `alert_skill`,
538-
`search_tool`, `general_question`, `guardrail`, `dashboard_summary`.
538+
`search_tool`, `general_question`, `guardrail`, `dashboard_summary`, and the agentic kinds
539+
(`agentic_obfuscation` is described below).
539540

540541
### `dashboard_summary` items
541542

@@ -574,6 +575,53 @@ The `expected_output` rubric:
574575
Each criterion is scored independently by the LLM judge, so `quality_score`
575576
is the fraction of satisfied criteria.
576577

578+
### `agentic_obfuscation` items
579+
580+
gen-ai masks sensitive values out of the conversation it stores and the trace it exports to
581+
Langfuse, while the model and the user's stream keep what was typed. So these items are not
582+
graded on the answer: they plant synthetic *canaries* in one or more user turns, then read back
583+
the stored conversation (`GET …/chat/conversations/{id}/items`) and every Langfuse trace of the
584+
session, and decide by exact substring. No LLM decides whether a value leaked.
585+
586+
```json
587+
{
588+
"id": "obfuscation-001",
589+
"dataset_name": "agent_obfuscation",
590+
"test_kind": "agentic_obfuscation",
591+
"question": ["My email is qa.canary5a1e@example.invalid. Show Total Sales by month.", "Now by quarter instead."],
592+
"expected_output": {
593+
"status": "enforced",
594+
"canaries": [
595+
{"nonce": "canary5a1e", "value": "qa.canary5a1e@example.invalid", "class": "EMAIL",
596+
"absent_from": ["conversation_db", "langfuse_trace"], "mask_marker_present": "[EMAIL]"}
597+
]
598+
}
599+
}
600+
```
601+
602+
- `question` is a string, or a list of turns sent in order to one conversation (from Langfuse, an
603+
input list). A question that is itself a JSON document must be stored in Langfuse as
604+
`{"query": "<the JSON>"}`: Langfuse parses a JSON-looking string input into an object.
605+
- A canary lists the sinks it must be `absent_from` and the sinks it must stay `present_in`. With
606+
`status: known_limitation` a `present_in` value is a known gap: the item fails once it closes, so
607+
the fixture is flipped deliberately. With `status: enforced` it guards a value that is not
608+
sensitive: masking it fails the item as `OVER-MASKED`. `record_only_paths` names sink paths
609+
reported but not gated.
610+
- `expected_turn_rejected: {"status_code": 422, "reason": "DATA_OBFUSCATION_CONTENT_REJECTED"}`
611+
expects the first turn to be refused.
612+
- `observe` (`automation_match`, `metric_title`, `stream_markers`) reports what the chat created
613+
or streamed, never gates it, and deletes the alert, export or metric it recognises.
614+
- Every turn needs a non-sensitive fragment of 12+ word characters, the *anchor*: a sink read back
615+
without every anchor is reported as blind, never as clean. `anchors` overrides the derived ones.
616+
- Langfuse is polled for the session's traces for 60 s (`GD_EVAL_OBFUSCATION_LANGFUSE_TIMEOUT_SEC`).
617+
When none arrives the run fails with `TRACE_NOT_FOUND`, the report's `trace_found` is false and
618+
`obfuscation_trace_found` is 0 -- a missing trace points at the export, never counts as a pass.
619+
620+
A leak in any run fails the item whatever `--gate` says. Before the first item of a workspace a
621+
preflight probe checks that both legs mask and both read-backs work, and stops every item of that
622+
workspace when they do not. The kind needs the `LANGFUSE_*` credentials of the Langfuse project
623+
the environment exports to: the traces are read, not only scored.
624+
577625
## Supported test kinds
578626

579627
| test_kind | What the agent must produce | Extra required |
@@ -585,6 +633,7 @@ is the fraction of satisfied criteria.
585633
| `general_question` | Text answer judged by LLM | `[llm-judge]` |
586634
| `guardrail` | Refusal/redirect (visualization response auto-fails) | `[llm-judge]` |
587635
| `dashboard_summary` | Dashboard summary (via `/summary` endpoint) scored against a rubric by LLM | `[llm-judge]` |
636+
| `agentic_obfuscation` | Canaries masked in the stored conversation and the Langfuse trace (exact match) | `LANGFUSE_*` |
588637

589638
## Optional extras
590639

@@ -625,6 +674,9 @@ the item's own root span. On the agentic path each score is mirrored onto the ag
625674
| `quality_score` | Fraction of strict check flags that are `True` (0.0–1.0). Shown in CLI as a percentage. |
626675
| `value_score` | Weighted blend: 0.6 × quality + 0.2 × speed (speed = max(0, 1 − latency/60s)). |
627676
| `latency_s` | Average per-run latency in seconds. |
677+
| `obfuscation_pass` / `obfuscation_no_leak` | `agentic_obfuscation`: the run's verdict, and whether no canary leaked. The comment names what failed, with canary values replaced by their class. |
678+
| `obfuscation_trace_found` | `agentic_obfuscation`: 0 when Langfuse held no trace for the conversation within the wait (`TRACE_NOT_FOUND`). |
679+
| `obfuscation_record_only_hit` | `agentic_obfuscation`: 1 when a record-only path held a canary; the comment carries every record-only observation. |
628680
| `provider_type` | Model vendor + gateway label (e.g. `ANTHROPIC`, `BEDROCK/ANTHROPIC`, `AZURE/OPENAI`). Stored in Langfuse trace metadata and tags. |
629681

630682
Score names carry no K; K and the gate are on the dataset-run metadata as `eval_k` and `eval_gate`.

‎packages/gooddata-eval/src/gooddata_eval/cli/agentic_runner.py‎

Lines changed: 18 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -18,6 +18,7 @@
1818
from gooddata_eval.core.agentic.guardrail import evaluate_agentic_guardrail
1919
from gooddata_eval.core.agentic.kda_skill import evaluate_agentic_kda_skill
2020
from gooddata_eval.core.agentic.metric_skill import evaluate_agentic_metric_skill
21+
from gooddata_eval.core.agentic.obfuscation import evaluate_agentic_obfuscation
2122
from gooddata_eval.core.agentic.search_tool import evaluate_agentic_search_tool
2223
from gooddata_eval.core.agentic.visualization import evaluate_agentic_visualization
2324
from gooddata_eval.core.agentic.what_if import evaluate_agentic_what_if
@@ -49,6 +50,7 @@ class _LfKw(TypedDict, total=False):
4950
"agentic_conversation",
5051
"agentic_kda_skill",
5152
"agentic_what_if",
53+
"agentic_obfuscation",
5254
}
5355
)
5456

@@ -88,7 +90,9 @@ class _LfKw(TypedDict, total=False):
8890
# metric skill. agentic_kda_skill is here on suspicion rather than proof: it triggers
8991
# create_key_driver_analysis with no cleanup, and while the evaluator only ever reads that
9092
# call's ARGUMENTS -- never a created object id -- whether the platform persists anything is
91-
# unverified. Move it to the allowlist once someone confirms it does not.
93+
# unverified. Move it to the allowlist once someone confirms it does not. agentic_obfuscation
94+
# items may ask for an alert, a scheduled export or a metric; it deletes what it recognises,
95+
# but that is cleanup, not read-only.
9296
#
9397
# agentic_dashboard_skill is absent by default rather than by evidence: gen-ai holds the draft and
9498
# any chart it authors in conversation state and writes neither until a user saves from the UI, so
@@ -278,6 +282,19 @@ def _dispatch_agentic(
278282
agent_id=agent_id,
279283
**lf_kw,
280284
)
285+
elif kind == "agentic_obfuscation":
286+
return evaluate_agentic_obfuscation(
287+
host=host,
288+
token=token,
289+
workspace_id=workspace_id,
290+
question=item.question,
291+
expected_output=eo if isinstance(eo, dict) else {},
292+
k=k,
293+
turns=item.turns,
294+
gate=gate,
295+
agent_id=agent_id,
296+
**lf_kw,
297+
)
281298
elif kind == "agentic_conversation":
282299
fixture_data = eo.get("fixture") or eo if isinstance(eo, dict) else {}
283300
return evaluate_agentic_conversation(

‎packages/gooddata-eval/src/gooddata_eval/cli/main.py‎

Lines changed: 12 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -66,6 +66,17 @@ def close(self) -> None:
6666
backend.close()
6767

6868

69+
def _positive_int(value: str) -> int:
70+
"""An argparse type for counts that must be at least 1."""
71+
try:
72+
number = int(value)
73+
except ValueError as exc:
74+
raise argparse.ArgumentTypeError(f"not an integer: {value!r}") from exc
75+
if number < 1:
76+
raise argparse.ArgumentTypeError(f"must be at least 1, got {number}")
77+
return number
78+
79+
6980
def _build_parser() -> argparse.ArgumentParser:
7081
parser = argparse.ArgumentParser(prog="gd-eval", description="Evaluate the GoodData AI agent.")
7182
sub = parser.add_subparsers(dest="command", required=True)
@@ -102,7 +113,7 @@ def _build_parser() -> argparse.ArgumentParser:
102113
"Default: workspace's current active model."
103114
),
104115
)
105-
run.add_argument("--runs", type=int, default=2, help="Independent runs per item. Default 2.")
116+
run.add_argument("--runs", type=_positive_int, default=2, help="Independent runs per item. Default 2.")
106117
run.add_argument(
107118
"--gate",
108119
choices=get_args(EvalGate),
Lines changed: 223 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,223 @@
1+
# (C) 2026 GoodData Corporation. All rights reserved.
2+
"""Deterministic leak verdict for the agentic_obfuscation kind.
3+
4+
Pure functions only: the caller collects the sinks, this module decides. A sink is any
5+
JSON-like document -- the stored conversation read back over the API, or every Langfuse
6+
trace of the conversation's session. A canary is found by exact substring over every string
7+
leaf, so the verdict never depends on an LLM. JSON serialised inside a string (Langfuse
8+
keeps span input that way) is decoded and walked too.
9+
10+
A sink that cannot be seen must not pass an absence check. Every turn therefore carries an
11+
anchor, a non-sensitive fragment of the question, and a sink that does not show every
12+
anchor is reported as blind instead of clean.
13+
"""
14+
15+
from __future__ import annotations
16+
17+
import json
18+
import re
19+
from collections.abc import Iterator, Mapping, Sequence
20+
from dataclasses import dataclass, field
21+
from typing import Any
22+
23+
SINK_CONVERSATION_DB = "conversation_db"
24+
SINK_LANGFUSE_TRACE = "langfuse_trace"
25+
SINKS = (SINK_CONVERSATION_DB, SINK_LANGFUSE_TRACE)
26+
27+
# Classes whose value may be written with separators the detector ignores, so a leak of
28+
# "4916 3385 0608 2832" must still match the compact nonce "4916338506082832".
29+
_DIGIT_CLASSES = frozenset({"CREDIT_CARD", "IBAN"})
30+
_SEPARATORS = re.compile(r"[\s\-]")
31+
# Placeholder the AI-on-data trace redaction writes over Langfuse input and output. Behind it
32+
# the Langfuse leg shows nothing, so an absence check there would prove nothing.
33+
AI_ON_DATA_REDACTION = "<redacted: AI-on-data enabled>"
34+
_ANCHOR_RUN = re.compile(r"[^\W_](?:[\w ]*[^\W_])?")
35+
_ANCHOR_MIN = 12
36+
_ANCHOR_MAX = 40
37+
38+
39+
@dataclass(frozen=True)
40+
class Occurrence:
41+
sink: str
42+
path: str
43+
needle: str
44+
45+
46+
@dataclass
47+
class CaseVerdict:
48+
failures: list[str] = field(default_factory=list)
49+
notes: list[str] = field(default_factory=list)
50+
51+
@property
52+
def passed(self) -> bool:
53+
return not self.failures
54+
55+
56+
def iter_strings(value: Any, path: str = "$") -> Iterator[tuple[str, str]]:
57+
"""Yield ``(path, text)`` for every string or number leaf and every key, decoding JSON held in a string.
58+
59+
A string that parses as a JSON document is walked instead of being yielded whole, so a
60+
value spelled with escapes (``\\u0040``) is seen decoded and reported once. Numbers are
61+
leaves too: a numeric secret in a pasted JSON document decodes to an int. Keys are text a
62+
user can type as well -- a pasted config may key on an email -- so each is yielded at the
63+
path of its object with a ``<key>`` suffix.
64+
"""
65+
if isinstance(value, bool):
66+
return
67+
if isinstance(value, int | float):
68+
yield path, str(value)
69+
return
70+
if isinstance(value, str):
71+
if value.lstrip()[:1] in ("{", "["):
72+
try:
73+
decoded = json.loads(value)
74+
except ValueError:
75+
pass
76+
else:
77+
yield from iter_strings(decoded, f"{path}<json>")
78+
return
79+
yield path, value
80+
elif isinstance(value, Mapping):
81+
for key, child in value.items():
82+
if isinstance(key, str):
83+
yield f"{path}.{key}<key>", key
84+
yield from iter_strings(child, f"{path}.{key}")
85+
elif isinstance(value, Sequence) and not isinstance(value, bytes | bytearray):
86+
for index, child in enumerate(value):
87+
yield from iter_strings(child, f"{path}[{index}]")
88+
89+
90+
def canary_needles(canary: Mapping[str, Any]) -> list[str]:
91+
"""Every spelling of the canary a leak may take: nonce, value and its declared variants."""
92+
keys = ("nonce", "value", "unescaped_value", "decoded_value", "compact_value")
93+
needles = [canary[key] for key in keys if isinstance(canary.get(key), str) and canary[key]]
94+
return list(dict.fromkeys(needles))
95+
96+
97+
def find_canary(canary: Mapping[str, Any], sink: str, document: Any) -> list[Occurrence]:
98+
needles = canary_needles(canary)
99+
compact = {_SEPARATORS.sub("", n) for n in needles} if canary.get("class") in _DIGIT_CLASSES else set()
100+
found: list[Occurrence] = []
101+
for path, text in iter_strings(document):
102+
hit = next((n for n in needles if n in text), None)
103+
if hit is None and compact:
104+
squeezed = _SEPARATORS.sub("", text)
105+
hit = next((n for n in compact if n in squeezed), None)
106+
if hit is not None:
107+
found.append(Occurrence(sink, path, hit))
108+
return found
109+
110+
111+
def contains(document: Any, needle: str) -> bool:
112+
return any(needle in text for _, text in iter_strings(document))
113+
114+
115+
def derive_anchor(question: str, canaries: Sequence[Mapping[str, Any]]) -> str | None:
116+
"""The longest plain-word run of the question once every canary spelling is cut out.
117+
118+
Word characters and spaces only, so escaping, JSON quoting and masking next to a canary
119+
cannot alter it between what was sent and what a sink stores.
120+
"""
121+
fragments = [question]
122+
# Longest spelling first: cutting the nonce out of "john\\.canary2d8e@..." first would
123+
# leave "john\\." behind, and the anchor would then end in text the mask replaces.
124+
needles = sorted({n for canary in canaries for n in canary_needles(canary)}, key=len, reverse=True)
125+
for needle in needles:
126+
fragments = [piece for fragment in fragments for piece in fragment.split(needle)]
127+
runs = [match.group(0) for fragment in fragments for match in _ANCHOR_RUN.finditer(fragment)]
128+
best = max(runs, key=len, default="")
129+
if len(best) < _ANCHOR_MIN:
130+
return None
131+
return best[:_ANCHOR_MAX].rstrip()
132+
133+
134+
def _declared_sinks(canary: Mapping[str, Any], key: str) -> list[str]:
135+
sinks = canary.get(key) or []
136+
unknown = [s for s in sinks if s not in SINKS]
137+
if unknown:
138+
raise ValueError(f"canary {canary.get('nonce')!r} names unknown sink(s) {unknown} in {key}")
139+
return list(sinks)
140+
141+
142+
def _format(occurrences: Sequence[Occurrence], limit: int = 5) -> str:
143+
shown = ", ".join(f"{o.path} ({o.needle!r})" for o in occurrences[:limit])
144+
more = len(occurrences) - limit
145+
return shown + (f" and {more} more" if more > 0 else "")
146+
147+
148+
def evaluate_case(
149+
expected: Mapping[str, Any],
150+
sinks: Mapping[str, Any],
151+
anchors: Sequence[str],
152+
*,
153+
turn_rejected: bool = False,
154+
) -> CaseVerdict:
155+
"""Decide one item from the collected sinks.
156+
157+
``sinks`` maps a sink name to its document, or to ``None`` when it was not collected
158+
(only allowed for the Langfuse leg of a rejected turn, which may export nothing).
159+
"""
160+
verdict = CaseVerdict()
161+
status = expected.get("status", "enforced")
162+
visible: dict[str, bool] = {}
163+
164+
for sink in SINKS:
165+
document = sinks.get(sink)
166+
if document is None:
167+
visible[sink] = False
168+
if not turn_rejected:
169+
verdict.failures.append(f"{sink}: not collected, so no absence claim can be made")
170+
continue
171+
if sink == SINK_LANGFUSE_TRACE and contains(document, AI_ON_DATA_REDACTION):
172+
verdict.failures.append(
173+
f"{sink}: trace input/output is replaced by {AI_ON_DATA_REDACTION!r} (enableAiOnData with "
174+
"enableGenAiTraceRedaction), so the obfuscation leg cannot be observed here"
175+
)
176+
visible[sink] = False
177+
continue
178+
missing = [anchor for anchor in anchors if not contains(document, anchor)]
179+
if missing and not turn_rejected:
180+
verdict.failures.append(f"{sink}: blind -- anchor(s) {missing} not found, the read-back is incomplete")
181+
visible[sink] = False
182+
continue
183+
visible[sink] = True
184+
185+
for canary in expected.get("canaries", []):
186+
label = f"{canary.get('class')} {canary.get('nonce')!r}"
187+
for sink in _declared_sinks(canary, "absent_from"):
188+
# A blind sink still convicts: a canary it does show is a leak all the same.
189+
if sinks.get(sink) is None:
190+
continue
191+
occurrences = find_canary(canary, sink, sinks[sink])
192+
# Paths the item declares as not yet decided (e.g. conversation state the SC does
193+
# not rule on) are reported, never gated, until a decision turns them into leaks.
194+
record_only = [re.compile(p) for p in canary.get("record_only_paths") or []]
195+
recorded = [o for o in occurrences if any(r.search(o.path) for r in record_only)]
196+
gated = [o for o in occurrences if o not in recorded]
197+
if gated:
198+
verdict.failures.append(f"LEAK {label} in {sink}: {_format(gated)}")
199+
if recorded:
200+
verdict.notes.append(f"RECORDED {label} in {sink} (record-only path, not gated): {_format(recorded)}")
201+
marker = canary.get("mask_marker_present")
202+
if marker and not turn_rejected and visible.get(sink) and not contains(sinks[sink], marker):
203+
verdict.failures.append(f"{label}: mask marker {marker!r} missing from {sink}")
204+
for sink in _declared_sinks(canary, "present_in"):
205+
if not visible.get(sink):
206+
continue
207+
if not find_canary(canary, sink, sinks[sink]):
208+
if status == "known_limitation":
209+
flip = (expected.get("known_limitation") or {}).get("flip_when", "")
210+
verdict.failures.append(
211+
f"{label} expected in {sink} ({status}) but is now masked -- the limitation no longer "
212+
f"reproduces, update the fixture deliberately. flip_when: {flip}"
213+
)
214+
else:
215+
# An enforced present_in guards a value that is not sensitive: masking it is
216+
# the defect (a false positive), not progress.
217+
verdict.failures.append(
218+
f"OVER-MASKED {label} in {sink}: expected unchanged but it was masked (false positive)"
219+
)
220+
221+
if turn_rejected and sinks.get(SINK_LANGFUSE_TRACE) is None:
222+
verdict.notes.append("rejected turn exported no Langfuse trace; only the database leg was asserted")
223+
return verdict

0 commit comments

Comments
 (0)