Skip to content

[DRAFT] Document data sequences, capture on demand, and sequence datasets - #5374

Open
Eliza Farley (elizafarley) wants to merge 28 commits into
mainfrom
sequences-docs
Open

Eliza Farley (elizafarley) wants to merge 28 commits into
mainfrom
sequences-docs

Conversation

@elizafarley

Copy link
Copy Markdown
Contributor

Summary

Documents data sequences end to end: capturing on demand, recording and reviewing sequences, building sequence datasets, and exporting them for custom training.

New pages

  • data/capture-sync/capture-on-demand.md: start and stop capture on demand with the capture-control module (default path), with writing your own capture control sensor as an advanced option.
  • data/sequences.md: what sequences are, and how to record, view, and manage them.
  • data/sequences-tutorial.md: end-to-end tutorial from recording to export, with Viam app and CLI/SDK tabs for each step.
  • train/sequence-dataset-format.md: the Parquet export format for sequence datasets and how to train on it.
  • Glossary term: sequence.

Reworked pages

  • train/create-a-dataset.md: organized around the two dataset types (image and sequence). One shared create step, then adding data, image training prep, and export split by type, plus a comparison table of what each type supports. Incoming links that pointed at removed section anchors now point at the page itself.
  • Data, train, CLI, and API reference pages updated where sequences or sequence datasets affect them: custom training scripts, train a model, export data, the CLI dataset commands, the ML training client, and others.

Generated SDK docs

  • Added sdk_protos_map.csv rows and proto description overrides for the sequence and sequence-dataset-export RPCs.

Checks

  • prettier 3.2.5, markdownlint, vale, and make build-prod pass on the changed pages.

🤖 Generated with Claude Code

Add how-tos for controlling capture with a sensor, recording and viewing
sequences, and creating a sequence dataset, plus a reference for the
sequence dataset export format. Update the data, train, and CLI pages
they affect, add the sequence glossary term, and add generator rows and
descriptions for the sequence and export RPCs.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Clarify the sequence and capture control sensor pages, fold the sequence
dataset into the dataset page with a type chooser, and add an end-to-end
tutorial.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Add a chooser between the module and a custom sensor, base the sequences
tutorial on the module, and document the module's commands and behavior.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Match UI labels for the capture-control and data manager blocks, the
sequence review details, and the dataset Add data button. Drop the
cleanup step.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uency

Each tutorial step now has a Viam app tab and a CLI and SDK tab, so the
whole tutorial can run from a terminal or an AI coding agent. Tested end
to end on a throwaway machine.

The capture-control module sends default_capture_frequency_hz from
startup until the first start_capture, so setting it to 2 captured
continuously before any recording. Leave it at 0 and pass frequency_hz
in start_capture instead, and describe the attribute accurately.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lated pages

Move the CLI and SDK setup into an expand and rewrite its lead-in.
Link the sequences tutorial from the capture control sensor, sequence
dataset, dataset format, and custom training script pages. Rename the
data management tutorial's sidebar title from Tutorial to Capture and
query tutorial so it's distinct from the sequences tutorial. Use the
Add data button label for empty datasets.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rename the page to "Start and stop capture on demand" (capture-on-demand.md)
and frame it around the job: recording several components together during
specific moments, and recording sequences. Make the capture-control module
the main path, move writing your own sensor to an Advanced section at the
bottom, and move the page to weight 7, after Start and Stop data capture.
Update inbound links.

Also includes in-progress edits to the sequences, glossary, and
create-a-dataset pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Share one create step for image and sequence datasets, then split
adding data, image training prep, and export by type. Point incoming
links at the page instead of removed section anchors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@netlify

netlify Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for viam-docs ready!

Name Link
🔨 Latest commit 75bb2aa
🔍 Latest deploy log https://app.netlify.com/projects/viam-docs/deploys/6ac595640570e40008cbd050
😎 Deploy Preview https://deploy-preview-5374--viam-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.
Lighthouse
Lighthouse
1 paths audited
Performance: 36 (🔴 down 7 from production)
Accessibility: 100 (no change from production)
Best Practices: 100 (no change from production)
SEO: 92 (no change from production)
PWA: 60 (no change from production)
View the detailed breakdown and full score reports

To edit notification comments on pull requests, go to your Netlify project configuration.

@viambot viambot added the safe to build This pull request is marked safe to build from a trusted zone label Sep 29, 2026
@elizafarley Eliza Farley (elizafarley) changed the title Document data sequences, capture on demand, and sequence datasets [DRAFT] Document data sequences, capture on demand, and sequence datasets Sep 29, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Tested on viam-server 1.9.0: without depends_on, the data manager logs one
startup error about the capture control sensor, then picks the sensor up and
captures normally. The setting isn't needed, so remove it from the sequences
tutorial, capture on demand, and the data reference. Also includes the
in-progress capture on demand rewrite.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The CLI can't add a registry module entry, so the sensor step now points to
the Viam MCP server. With depends_on gone, the data manager step works from
the CLI using add-resource --api and resource update. Remove the script and
its JSON files from the tutorial and capture on demand.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Rename "Image dataset" to "Binary dataset" and note per-SDK dataset-type
support in the Create a dataset tabs. Drop the "control capture with a
sensor" approach from Filter at the edge and remove now-redundant
capture-on-demand links from conditional sync and retrain.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fix incorrect method names, CLI steps, and UI steps; document
capture control sensor edge cases, SDK coverage, and export paths;
align dataset terminology and tab names; flag open questions for
engineering as TODO comments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rial

The Python SDK now has create_sequence, get_sequence, update_sequence,
delete_sequence, and list_sequences (viam-python-sdk#1277).

- Tutorial: replace the TypeScript list_sequences script with Python,
  drop the Node.js prerequisite, rename the CLI and SDK tabs to CLI and Python.
- Sequences page: use SequenceResourceFilter in the create example, remove
  statements that Python lacks these methods, note that update_sequence
  builds the field mask automatically.
- sdk_protos_map.csv: add Python method names for the five RPCs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0199tsKsv8WPJWQuMoa4M8q8
@jeremyrose-viam

Copy link
Copy Markdown
Member

FYI, the Go DataClient sequences rows here also cover the sequences methods that #5190 tracks as undocumented, so that issue can close when this merges. The Go scraper already attributes by proto name, so no generator change was needed.

- Remove the per-step Viam app and CLI/SDK tab walkthroughs from
  capture-on-demand, leaving the config and DoCommand reference
- Reword the sequences tutorial prerequisites for CLI and Python 3
- Rename the dataset-type column from "Image" to "Binary dataset"
- Reword the custom training script intro and link the sequence
  dataset format from the dataset-reading notes
- Broaden the Hugo shortcode-ignore vale regex to cover closing and
  hyphenated shortcode tags

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Scope custom-training-scripts steps, arguments, export, and local testing
to binary or sequence datasets, add a direct-run command for sequence
scripts, and use the app's term "binary dataset" instead of "image
dataset" in create-a-dataset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L1Us1m2SEqcV181b49UFLt
Explain how to connect a data client, find the part ID, find capture
times, and that the start and end times are inclusive.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread docs/train/create-a-dataset.md Outdated
- Managed [training](/train/train-a-model/) doesn't accept sequence datasets.
Train with a custom training script.
- Sequence exports run queries against your organization's data, and tabular queries count toward your data query usage.
<!-- TODO(eng): confirm whether tabular queries in a sequence export count toward data query usage. No usage accounting found in the export path. -->

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We query ADF to get tabular data and we bill for ADF, I think at the end of the month.

- Correct export timestamp to capture time
- Unparseable capture control readings close sequences without reverting capture
- Call out view in data gallery/query page buttons on a sequence
- Clarify deleting sequences vs. deleting data
- Note CreateSequence needs only part access; other sequence APIs need org keys
- Note sequences are scoped to one part, remotes included
- Resolve eng TODO: tabular sequence exports are billed as data queries
- Regenerate data and dataset API pages with sequence methods

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@elizafarley
Eliza Farley (elizafarley) marked this pull request as ready for review October 5, 2026 21:32
The sequence CRUD methods (create/get/update/delete/list_sequences)
first shipped in viam-sdk 0.83.1; 0.83.0 doesn't have them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The failure section described four bad-readings cases that module users
can't trigger, and two of them didn't match rdk's poller. Custom sensor
authors are covered by the readings reference.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Tested on a throwaway machine (viam-server 1.9.0, capture-control 0.1.1)
with a fake camera on a remote named rmt: resource_name "cam" captured
and synced as component_name "cam", and the sequence returned those
images; "rmt:cam" was rejected as an unknown resource on every poll.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
create_sequence, get_sequence, update_sequence, delete_sequence, and
list_sequences are now published on python.viam.dev, so the generator
picks up their Python tabs.

Also revert the unrelated Vale Hugo TokenIgnores/BlockIgnores change
to match main.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015CiwX7xSPujt2kG9kbAVjE

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed from a documentation-user perspective, and I ran the tutorial's CLI path end to end on a fresh machine. Steps 1 to 7 work as written, and the exported Parquet schema matches the format page. Confirmed by running: the CLI can't add the module (step 2), and --api is required for the data manager (step 3). Comments are inline. The ones I'd prioritize: the stale capture_disabled row, the version prerequisites, the unresolved billing wording, and the missing sequence training example.

Not in the diff: the contradiction with cli/configure-machines.md on --api (see the comment on tutorial line 189).

Comment thread docs/data/reference.md Outdated
Comment thread docs/data/reference.md Outdated
Comment thread docs/data/sequences-tutorial.md
Comment thread docs/data/sequences-tutorial.md
Comment thread docs/data/sequences-tutorial.md Outdated
Comment thread docs/train/create-a-dataset.md
Comment thread docs/train/create-a-dataset.md Outdated
Comment thread docs/train/sequence-dataset-format.md
Comment thread docs/train/sequence-dataset-format.md Outdated
Comment thread docs/train/custom-training-scripts.md
- Name the non-sequence dataset type "binary" everywhere, matching the UI
- Link sequence export usage note to the billing page
- Clarify capture control tags replace (not merge with) the data manager's tags
- Recommend the latest viam-server, listing per-feature minimum versions
- Note capture control sensors can enable capture in the capture_disabled row
- Document that non-JPEG/PNG images are silently omitted from sequences
- Document ListSequences pagination, payload shape, and UTC timestamps
- Tutorial: fix sequence length and image counts, show full columns, note
  the export steps run in a terminal from one directory

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@jeremyrose-viam

Copy link
Copy Markdown
Member

Thanks for the quick turnaround on the review. The tutorial now reads cleanly from start to finish. I'm approving, and I wanted to make a case for one thing that I'd really like to see, either in this PR or as a fast follow-up.

A short sequence-dataset example in custom-training-scripts.md. Everything in this PR leads to "train a model on a sequence dataset," and the tutorial ends with the reader holding three Parquet files. The binary path on that page has a skeleton script to copy. The sequence path has an argument table and a pointer to the column docs, so the reader has to work out the same three things from scratch:

  • path is relative in an export and absolute in a cloud job.
  • payload is a JSON string with a readings wrapper.
  • Images and readings run at different rates, so they have to be matched by nearest timestamp. A row-by-row join gives wrong pairs without any error.

I don't mean a training loop, which depends on the task. About 15 lines that load and pair the data would do it. I ran this sketch against an export from the tutorial (87 images, 87 readings, all image paths resolved), so you're welcome to adapt it:

import json
import os

import pandas as pd

export_dir = "demos"  # in a training job, `path` is already absolute
images = pd.read_parquet(f"{export_dir}/parquet/binary_data.parquet")
readings = pd.read_parquet(f"{export_dir}/parquet/tabular_data.parquet")
sequences = pd.read_parquet(f"{export_dir}/parquet/sequences.parquet")

# Resolve image paths. They're relative in an export, absolute in a training job.
images["path"] = images["path"].map(
    lambda p: p if os.path.isabs(p) else os.path.join(export_dir, p)
)

# Unpack the JSON payload, for example {"readings": {"a": 1.0}}.
readings = pd.concat(
    [readings, pd.json_normalize(readings["payload"].map(json.loads))], axis=1
)

# Each resource has its own sample rate, so pair each image with the
# nearest reading in the same sequence.
pairs = pd.merge_asof(
    images.sort_values("timestamp"),
    readings.sort_values("timestamp"),
    on="timestamp",
    by="sequence_id",
    direction="nearest",
    suffixes=("_image", "_reading"),
).merge(sequences[["sequence_id", "tags"]], on="sequence_id")

Two things to confirm before using it. A training job reads --binary_data_file and the other arguments instead of hardcoding paths, and pd.merge_asof assumes pandas is available in the training container.

If you'd rather keep this PR small, a follow-up is fine. I'd just like it to land soon, because this is the first thing a reader will try after the tutorial.

@viam-overwatch

viam-overwatch Bot commented Oct 9, 2026

Copy link
Copy Markdown

Hey Eliza Farley (@elizafarley) — this PR has been approved and CI has been green for 3+ business days. Ready to merge?

Auto-comment from overwatch. Will not re-nudge for 7 days.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

safe to build This pull request is marked safe to build from a trusted zone

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants