Most robotics teams can tell you how many demonstrations they have collected. Far fewer can tell you what is actually inside one. Teleoperation demonstration logs are the unit of currency in robot learning, and a log missing three fields can make ten thousand episodes untrainable.
This guide walks through the four layers of a log, every field worth carrying and why, how intervention logs differ from demonstration logs, and the five schema decisions you cannot reverse once collection starts.
What a demonstration log is
A demonstration log is the complete record of one episode: everything the robot sensed, everything the operator commanded, and everything a human later said about how it went. It is not a video file with a name. It is a structured object, and how well it is structured determines whether the episode is trainable, auditable, and reusable two years from now.
Most teams design this schema after they have collected ten thousand episodes. That is the wrong order and it is expensive to correct.
The four layers of a log
Layer 1: observation
What the robot’s sensors recorded. Camera streams, depth, proprioceptive joint state, and any external sensing. Each stream needs its own timestamp series rather than an assumed constant rate, because dropped frames are normal and silent.
Layer 2: action
What was commanded. On a leader-follower rig this is the leader’s joint positions. On a headset system it is the retargeted end-effector pose. The critical detail is stating which convention you used and never changing it mid-program.
Layer 3: episode metadata
Who, what, when, and under what conditions. This is the layer teams skimp on and later wish they had.
Layer 4: outcome and annotation
Whether it worked, where the phase boundaries sit, and what went wrong. Some of this is automatic; the rest comes from annotation and labeling.
Field by field: what every log should carry
| Field | Layer | Why it matters later |
|---|---|---|
| episode_id | Metadata | Globally unique; never reuse after a deletion |
| session_id | Metadata | Groups episodes that share a calibration and a shift |
| operator_id | Metadata | Pseudonymous; lets you trace a defect to its source |
| rig_id and firmware | Metadata | Isolates hardware-specific artifacts across a fleet |
| calibration_ref | Metadata | Points at the intrinsics and extrinsics in force that session |
| task_id and variant | Metadata | Distinguishes the task from the specific staging of it |
| object_set | Metadata | Instance-level, not category-level; “red mug 04” not “mug” |
| scene_conditions | Metadata | Lighting, clutter level, surface; enables diversity audits |
| timestamps per stream | Observation | Absolute and monotonic; the foundation of everything |
| action convention | Action | Joint space or task space, absolute or delta, frame of reference |
| control_rate_actual | Action | Measured, not configured; drops reveal system stress |
| latency_stats | Action | Median, p95, worst spike. See latency budgets |
| gripper state | Action | Continuous where possible; binary loses grasp modulation |
| force_torque | Observation | The only honest record of contact quality |
| outcome | Outcome | Success, partial, failure, aborted; never just a boolean |
| failure_mode | Outcome | Controlled vocabulary, not free text |
| retry_count | Outcome | Recovery behavior is high-value training signal |
| phase_boundaries | Annotation | Reach, grasp, transport, place, release |
| qa_status and reviewer | Annotation | Auditable review trail for enterprise buyers |
| consent_ref | Metadata | Links to the applicable consent record |
Intervention logs are a different animal
Teleoperation logs assume a human drove the whole episode. Remote operations produce a different record entirely: the robot was autonomous, something went wrong, a human stepped in, and then autonomy resumed.
We run this kind of monitoring for an autonomous mobility company operating sidewalk delivery robots and self-driving passenger vehicles, and the log needs fields a demonstration log never has.
- trigger_type – what caused the handover: the robot asked, a monitor noticed, or a rule fired
- time_to_acknowledge – how long before a human engaged
- autonomy_state_at_handover – what the robot believed at the moment it stopped coping
- resolution_type – guided, remotely driven, escalated to field, or resolved itself
- time_to_resolve and time_to_resume_autonomy
- root_cause – controlled vocabulary: perception, planning, hardware, environment, third party
These fields are what turn an operations cost centre into a training data source, which is the subject of closing the loop from intervention to training data.
Five schema decisions you cannot easily reverse
- Action convention. Changing from joint space to end-effector space mid-program splits your dataset in two.
- Timestamp source. Pick one clock authority and one epoch. Mixed sources produce drift nobody notices for months.
- Failure vocabulary. Free-text failure notes cannot be aggregated. Define the list early and version it.
- Object identity granularity. Instance-level identifiers cost nothing at capture and are impossible to recover later.
- Storage layout. One episode per directory with a manifest survives migrations. Monolithic archives do not.
If you are evaluating a data partner, ask to see their log schema before you ask about price. It tells you more about the quality of their operation than any deck will. Our operator quality guide covers what else to ask.
A settled log schema is part of what makes a fast ramp possible. Our mobile manipulation case study covers a 90-day path from cold-start to production.
Frequently asked questions
Should failed episodes be logged the same way as successes?
Yes, identically, with the outcome and failure mode recorded. Failures are how a policy learns recovery, and a dataset of only successes teaches a robot that nothing ever goes wrong.
How much storage should we plan for?
Multi-camera capture grows faster than almost every first forecast. Plan for the full raw stream during active collection and a compressed archive tier afterwards, and decide the retention policy before volume forces the decision for you.
Do we need operator identity in the log?
You need a stable pseudonymous identifier so you can trace defects and measure per-operator quality. You do not need personal identity in the dataset itself, and keeping it out simplifies your privacy position considerably.
Can we standardize on an existing format?
Adopting a community format helps interoperability and makes pooling with public data easier. Most teams extend one rather than adopting it unchanged, because the metadata layer is where their own requirements live.
The log schema is the cheapest thing to get right at the start and the most expensive thing to retrofit. If you are designing one, or inherited one you do not trust, send us the shape of it and we will tell you what will hurt at scale.






