Skip to main content Diffraction

The collection field guide / 01

How to write a robotics dataset brief.

A practical guide to turning a model objective into tasks, signals and acceptance criteria a collection team can execute.

By Diffraction · September 8, 2026

From question to capture
01Define the task
02Choose the evidence
03Agree on acceptance
Download the brief template

Editable Markdown · No signup required

Decision 01

Start with the failure you need to fix.

Describe the behavior your system struggles with and how you will evaluate improvement. A collection for recognizing human actions has different requirements from one for learning robot control. State whether you need human observations, measured robot demonstrations, or both.

Worked example · Illustrative

Instead of ‘more kitchen video’: ‘Evaluate whether our action model can distinguish a reach, grasp and placement when hands partially occlude a small object.’

Put in the brief:Write one objective, one intended use and one downstream evaluation. Video of a person does not supply measured robot commands.

Decision 02

Make an episode unambiguous.

Define the starting state, instruction, end condition and allowed variations. Decide how to label an unsuccessful attempt, a recovery, an interruption and a reset. Give operators examples of each so the same instruction produces comparable records across sessions.

Worked example · Illustrative

For an object-transfer task: begin with the object inside a marked source area; end after release in the destination. Record a dropped object as a failed attempt, and label any retry separately.

Put in the brief:Count accepted episodes and coverage by condition alongside hours. Agree on which idle or setup periods count toward delivery.

Decision 03

Specify what must be measured.

List the views and sensor channels the evaluation needs. For each, define units, coordinate frames, timestamps, calibration and missing-data behavior. Identify the channels that may be estimated, including their processing version and confidence fields. A plausible overlay alone cannot establish accuracy.

Worked example · Illustrative

If depth is essential, request depth coverage and invalid-pixel behavior at the surfaces you care about. If a wrist trajectory is derived from images and depth, label it as an estimate and specify how it will be evaluated.

Put in the brief:Agree on synchronization tolerance and how residual error will be measured. Device IMU and wrist-mounted IMU describe different motion.

Decision 04

Design variation around deployment.

Describe the environments, participants, objects, lighting and occlusions the model should encounter. A controlled room makes setup repeatable; distributed collection can broaden environments. Neither guarantees useful diversity without an explicit coverage plan.

Worked example · Illustrative

Specify quotas for bright and dim settings, object sizes and occlusion conditions. Keep participants or sessions grouped when making evaluation splits so closely related frames do not leak across them.

Put in the brief:Ask for a manifest that lets you audit coverage. State what is known about participant and session identity, and how unknown identity is handled.

Decision 05

Turn quality into a pass/fail decision.

Define a metric, threshold, test method and review scope for each critical requirement. File integrity, signal coverage, annotation accuracy and downstream performance answer different questions. Decide which are required for pilot acceptance and who signs off.

Worked example · Illustrative

An integrity check can establish that a video decodes. It cannot establish that a hand label is correct. Use a separately reviewed subset to evaluate labels, then run your own task evaluation to assess model usefulness.

Put in the brief:Set project-specific thresholds during the pilot. Include representative hard cases and agree on rework or rejection before scaling.

Decision 06

Test the handoff before scaling.

Request an example package and run it through your loader. Agree on field names, shapes, units, schema version and null handling, along with native media, derived annotations, provenance and checksums. Review contributor permissions, intended uses and license terms with the responsible people before collection.

Worked example · Illustrative

Read an episode from first frame to last, inspect missing channels, map timestamps and reproduce a quality check. Confirm that the delivery contains the signals your training code actually consumes.

Put in the brief:Choose the format around the data and consumer. Export compatibility does not establish annotation quality or suitability for a particular policy.

Before the full collection

A pilot review you can use.

Use these checks to structure the conversation. Choose thresholds with your team; this is a planning checklist, not a universal quality standard.

Pilot acceptance criteria and evidence
CheckAcceptance questionEvidence to request
Video integrityEvery delivered file decodes; frame and manifest counts reconcile.Automated full-file checks plus a documented discrepancy report.
Task visibilityThe required interaction is visible for an agreed share of task frames.Review a defined sample, including occlusion and motion cases.
Signal alignmentRequired streams stay within the agreed timing tolerance.State the synchronization method and test residual error.
Annotation qualitySelected labels meet the agreed error criterion.Evaluate against a separately reviewed reference subset.
Consumer handoffThe buyer’s loader reads the agreed channels and missing values.Run an example episode in the target environment before scale-up.

See the distinction in a real sample.

Our kitchen release separates native RGB-D capture from estimated hand and object annotations. Its quality notes show why a working reader and intact files are only part of evaluation.

Inspect the egocentric kitchen sample ↗

Format references

Consult the format’s own documentation when agreeing on the delivery schema and consumer version.

Make the brief concrete

Bring us the hard part.

Share the task, the signal you need and the decision you want the pilot to answer. We’ll scope the capture with you.

Discuss your dataset brief ↗