Repository navigation
[DRAFT] Document data sequences, capture on demand, and sequence datasets - #5374
Eliza Farley (elizafarley) wants to merge 28 commits into
Conversation
Add how-tos for controlling capture with a sensor, recording and viewing sequences, and creating a sequence dataset, plus a reference for the sequence dataset export format. Update the data, train, and CLI pages they affect, add the sequence glossary term, and add generator rows and descriptions for the sequence and export RPCs. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Clarify the sequence and capture control sensor pages, fold the sequence dataset into the dataset page with a type chooser, and add an end-to-end tutorial. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Add a chooser between the module and a custom sensor, base the sequences tutorial on the module, and document the module's commands and behavior. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Match UI labels for the capture-control and data manager blocks, the sequence review details, and the dataset Add data button. Drop the cleanup step. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uency Each tutorial step now has a Viam app tab and a CLI and SDK tab, so the whole tutorial can run from a terminal or an AI coding agent. Tested end to end on a throwaway machine. The capture-control module sends default_capture_frequency_hz from startup until the first start_capture, so setting it to 2 captured continuously before any recording. Leave it at 0 and pass frequency_hz in start_capture instead, and describe the attribute accurately. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lated pages Move the CLI and SDK setup into an expand and rewrite its lead-in. Link the sequences tutorial from the capture control sensor, sequence dataset, dataset format, and custom training script pages. Rename the data management tutorial's sidebar title from Tutorial to Capture and query tutorial so it's distinct from the sequences tutorial. Use the Add data button label for empty datasets. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rename the page to "Start and stop capture on demand" (capture-on-demand.md) and frame it around the job: recording several components together during specific moments, and recording sequences. Make the capture-control module the main path, move writing your own sensor to an Advanced section at the bottom, and move the page to weight 7, after Start and Stop data capture. Update inbound links. Also includes in-progress edits to the sequences, glossary, and create-a-dataset pages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Share one create step for image and sequence datasets, then split adding data, image training prep, and export by type. Point incoming links at the page instead of removed section anchors. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
✅ Deploy Preview for viam-docs ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Tested on viam-server 1.9.0: without depends_on, the data manager logs one startup error about the capture control sensor, then picks the sensor up and captures normally. The setting isn't needed, so remove it from the sequences tutorial, capture on demand, and the data reference. Also includes the in-progress capture on demand rewrite. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The CLI can't add a registry module entry, so the sensor step now points to the Viam MCP server. With depends_on gone, the data manager step works from the CLI using add-resource --api and resource update. Remove the script and its JSON files from the tutorial and capture on demand. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Rename "Image dataset" to "Binary dataset" and note per-SDK dataset-type support in the Create a dataset tabs. Drop the "control capture with a sensor" approach from Filter at the edge and remove now-redundant capture-on-demand links from conditional sync and retrain. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fix incorrect method names, CLI steps, and UI steps; document capture control sensor edge cases, SDK coverage, and export paths; align dataset terminology and tab names; flag open questions for engineering as TODO comments. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rial The Python SDK now has create_sequence, get_sequence, update_sequence, delete_sequence, and list_sequences (viam-python-sdk#1277). - Tutorial: replace the TypeScript list_sequences script with Python, drop the Node.js prerequisite, rename the CLI and SDK tabs to CLI and Python. - Sequences page: use SequenceResourceFilter in the create example, remove statements that Python lacks these methods, note that update_sequence builds the field mask automatically. - sdk_protos_map.csv: add Python method names for the five RPCs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0199tsKsv8WPJWQuMoa4M8q8
|
FYI, the Go |
- Remove the per-step Viam app and CLI/SDK tab walkthroughs from capture-on-demand, leaving the config and DoCommand reference - Reword the sequences tutorial prerequisites for CLI and Python 3 - Rename the dataset-type column from "Image" to "Binary dataset" - Reword the custom training script intro and link the sequence dataset format from the dataset-reading notes - Broaden the Hugo shortcode-ignore vale regex to cover closing and hyphenated shortcode tags Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Scope custom-training-scripts steps, arguments, export, and local testing to binary or sequence datasets, add a direct-run command for sequence scripts, and use the app's term "binary dataset" instead of "image dataset" in create-a-dataset. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L1Us1m2SEqcV181b49UFLt
Explain how to connect a data client, find the part ID, find capture times, and that the start and end times are inclusive. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| - Managed [training](/train/train-a-model/) doesn't accept sequence datasets. | ||
| Train with a custom training script. | ||
| - Sequence exports run queries against your organization's data, and tabular queries count toward your data query usage. | ||
| <!-- TODO(eng): confirm whether tabular queries in a sequence export count toward data query usage. No usage accounting found in the export path. --> |
There was a problem hiding this comment.
We query ADF to get tabular data and we bill for ADF, I think at the end of the month.
- Correct export timestamp to capture time - Unparseable capture control readings close sequences without reverting capture - Call out view in data gallery/query page buttons on a sequence - Clarify deleting sequences vs. deleting data - Note CreateSequence needs only part access; other sequence APIs need org keys - Note sequences are scoped to one part, remotes included - Resolve eng TODO: tabular sequence exports are billed as data queries - Regenerate data and dataset API pages with sequence methods Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sequence CRUD methods (create/get/update/delete/list_sequences) first shipped in viam-sdk 0.83.1; 0.83.0 doesn't have them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The failure section described four bad-readings cases that module users can't trigger, and two of them didn't match rdk's poller. Custom sensor authors are covered by the readings reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Tested on a throwaway machine (viam-server 1.9.0, capture-control 0.1.1) with a fake camera on a remote named rmt: resource_name "cam" captured and synced as component_name "cam", and the sequence returned those images; "rmt:cam" was rejected as an unknown resource on every poll. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
create_sequence, get_sequence, update_sequence, delete_sequence, and list_sequences are now published on python.viam.dev, so the generator picks up their Python tabs. Also revert the unrelated Vale Hugo TokenIgnores/BlockIgnores change to match main. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015CiwX7xSPujt2kG9kbAVjE
Jeremy Rose (jeremyrose-viam)
left a comment
There was a problem hiding this comment.
Reviewed from a documentation-user perspective, and I ran the tutorial's CLI path end to end on a fresh machine. Steps 1 to 7 work as written, and the exported Parquet schema matches the format page. Confirmed by running: the CLI can't add the module (step 2), and --api is required for the data manager (step 3). Comments are inline. The ones I'd prioritize: the stale capture_disabled row, the version prerequisites, the unresolved billing wording, and the missing sequence training example.
Not in the diff: the contradiction with cli/configure-machines.md on --api (see the comment on tutorial line 189).
- Name the non-sequence dataset type "binary" everywhere, matching the UI - Link sequence export usage note to the billing page - Clarify capture control tags replace (not merge with) the data manager's tags - Recommend the latest viam-server, listing per-feature minimum versions - Note capture control sensors can enable capture in the capture_disabled row - Document that non-JPEG/PNG images are silently omitted from sequences - Document ListSequences pagination, payload shape, and UTC timestamps - Tutorial: fix sequence length and image counts, show full columns, note the export steps run in a terminal from one directory Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Thanks for the quick turnaround on the review. The tutorial now reads cleanly from start to finish. I'm approving, and I wanted to make a case for one thing that I'd really like to see, either in this PR or as a fast follow-up. A short sequence-dataset example in custom-training-scripts.md. Everything in this PR leads to "train a model on a sequence dataset," and the tutorial ends with the reader holding three Parquet files. The binary path on that page has a skeleton script to copy. The sequence path has an argument table and a pointer to the column docs, so the reader has to work out the same three things from scratch:
I don't mean a training loop, which depends on the task. About 15 lines that load and pair the data would do it. I ran this sketch against an export from the tutorial (87 images, 87 readings, all image paths resolved), so you're welcome to adapt it: import json
import os
import pandas as pd
export_dir = "demos" # in a training job, `path` is already absolute
images = pd.read_parquet(f"{export_dir}/parquet/binary_data.parquet")
readings = pd.read_parquet(f"{export_dir}/parquet/tabular_data.parquet")
sequences = pd.read_parquet(f"{export_dir}/parquet/sequences.parquet")
# Resolve image paths. They're relative in an export, absolute in a training job.
images["path"] = images["path"].map(
lambda p: p if os.path.isabs(p) else os.path.join(export_dir, p)
)
# Unpack the JSON payload, for example {"readings": {"a": 1.0}}.
readings = pd.concat(
[readings, pd.json_normalize(readings["payload"].map(json.loads))], axis=1
)
# Each resource has its own sample rate, so pair each image with the
# nearest reading in the same sequence.
pairs = pd.merge_asof(
images.sort_values("timestamp"),
readings.sort_values("timestamp"),
on="timestamp",
by="sequence_id",
direction="nearest",
suffixes=("_image", "_reading"),
).merge(sequences[["sequence_id", "tags"]], on="sequence_id")Two things to confirm before using it. A training job reads If you'd rather keep this PR small, a follow-up is fine. I'd just like it to land soon, because this is the first thing a reader will try after the tutorial. |
|
Hey Eliza Farley (@elizafarley) — this PR has been approved and CI has been green for 3+ business days. Ready to merge? Auto-comment from overwatch. Will not re-nudge for 7 days. |

Summary
Documents data sequences end to end: capturing on demand, recording and reviewing sequences, building sequence datasets, and exporting them for custom training.
New pages
data/capture-sync/capture-on-demand.md: start and stop capture on demand with the capture-control module (default path), with writing your own capture control sensor as an advanced option.data/sequences.md: what sequences are, and how to record, view, and manage them.data/sequences-tutorial.md: end-to-end tutorial from recording to export, with Viam app and CLI/SDK tabs for each step.train/sequence-dataset-format.md: the Parquet export format for sequence datasets and how to train on it.sequence.Reworked pages
train/create-a-dataset.md: organized around the two dataset types (image and sequence). One shared create step, then adding data, image training prep, and export split by type, plus a comparison table of what each type supports. Incoming links that pointed at removed section anchors now point at the page itself.Generated SDK docs
sdk_protos_map.csvrows and proto description overrides for the sequence and sequence-dataset-export RPCs.Checks
make build-prodpass on the changed pages.🤖 Generated with Claude Code