Reading a Teleoperation Demonstration Log: What Every Field Means

teleoperation demonstration logs

Most robotics teams can tell you how many demonstrations they have collected. Far fewer can tell you what is actually inside one. Teleoperation demonstration logs are the unit of currency in robot learning, and a log missing three fields can make ten thousand episodes untrainable.

This guide walks through the four layers of a log, every field worth carrying and why, how intervention logs differ from demonstration logs, and the five schema decisions you cannot reverse once collection starts.

Table of contents

    What a demonstration log is

    A demonstration log is the complete record of one episode: everything the robot sensed, everything the operator commanded, and everything a human later said about how it went. It is not a video file with a name. It is a structured object, and how well it is structured determines whether the episode is trainable, auditable, and reusable two years from now.

    Most teams design this schema after they have collected ten thousand episodes. That is the wrong order and it is expensive to correct.

    The four layers of a log

    Layer 1: observation

    What the robot’s sensors recorded. Camera streams, depth, proprioceptive joint state, and any external sensing. Each stream needs its own timestamp series rather than an assumed constant rate, because dropped frames are normal and silent.

    Layer 2: action

    What was commanded. On a leader-follower rig this is the leader’s joint positions. On a headset system it is the retargeted end-effector pose. The critical detail is stating which convention you used and never changing it mid-program.

    Layer 3: episode metadata

    Who, what, when, and under what conditions. This is the layer teams skimp on and later wish they had.

    Layer 4: outcome and annotation

    Whether it worked, where the phase boundaries sit, and what went wrong. Some of this is automatic; the rest comes from annotation and labeling.

    Field by field: what every log should carry

    Field Layer Why it matters later
    episode_id Metadata Globally unique; never reuse after a deletion
    session_id Metadata Groups episodes that share a calibration and a shift
    operator_id Metadata Pseudonymous; lets you trace a defect to its source
    rig_id and firmware Metadata Isolates hardware-specific artifacts across a fleet
    calibration_ref Metadata Points at the intrinsics and extrinsics in force that session
    task_id and variant Metadata Distinguishes the task from the specific staging of it
    object_set Metadata Instance-level, not category-level; “red mug 04” not “mug”
    scene_conditions Metadata Lighting, clutter level, surface; enables diversity audits
    timestamps per stream Observation Absolute and monotonic; the foundation of everything
    action convention Action Joint space or task space, absolute or delta, frame of reference
    control_rate_actual Action Measured, not configured; drops reveal system stress
    latency_stats Action Median, p95, worst spike. See latency budgets
    gripper state Action Continuous where possible; binary loses grasp modulation
    force_torque Observation The only honest record of contact quality
    outcome Outcome Success, partial, failure, aborted; never just a boolean
    failure_mode Outcome Controlled vocabulary, not free text
    retry_count Outcome Recovery behavior is high-value training signal
    phase_boundaries Annotation Reach, grasp, transport, place, release
    qa_status and reviewer Annotation Auditable review trail for enterprise buyers
    consent_ref Metadata Links to the applicable consent record

    Intervention logs are a different animal

    Teleoperation logs assume a human drove the whole episode. Remote operations produce a different record entirely: the robot was autonomous, something went wrong, a human stepped in, and then autonomy resumed.

    We run this kind of monitoring for an autonomous mobility company operating sidewalk delivery robots and self-driving passenger vehicles, and the log needs fields a demonstration log never has.

    • trigger_type – what caused the handover: the robot asked, a monitor noticed, or a rule fired
    • time_to_acknowledge – how long before a human engaged
    • autonomy_state_at_handover – what the robot believed at the moment it stopped coping
    • resolution_type – guided, remotely driven, escalated to field, or resolved itself
    • time_to_resolve and time_to_resume_autonomy
    • root_cause – controlled vocabulary: perception, planning, hardware, environment, third party

    These fields are what turn an operations cost centre into a training data source, which is the subject of closing the loop from intervention to training data.

    Five schema decisions you cannot easily reverse

    1. Action convention. Changing from joint space to end-effector space mid-program splits your dataset in two.
    2. Timestamp source. Pick one clock authority and one epoch. Mixed sources produce drift nobody notices for months.
    3. Failure vocabulary. Free-text failure notes cannot be aggregated. Define the list early and version it.
    4. Object identity granularity. Instance-level identifiers cost nothing at capture and are impossible to recover later.
    5. Storage layout. One episode per directory with a manifest survives migrations. Monolithic archives do not.

    If you are evaluating a data partner, ask to see their log schema before you ask about price. It tells you more about the quality of their operation than any deck will. Our operator quality guide covers what else to ask.

    A settled log schema is part of what makes a fast ramp possible. Our mobile manipulation case study covers a 90-day path from cold-start to production.

    Frequently asked questions

    Should failed episodes be logged the same way as successes?

    Yes, identically, with the outcome and failure mode recorded. Failures are how a policy learns recovery, and a dataset of only successes teaches a robot that nothing ever goes wrong.

    How much storage should we plan for?

    Multi-camera capture grows faster than almost every first forecast. Plan for the full raw stream during active collection and a compressed archive tier afterwards, and decide the retention policy before volume forces the decision for you.

    Do we need operator identity in the log?

    You need a stable pseudonymous identifier so you can trace defects and measure per-operator quality. You do not need personal identity in the dataset itself, and keeping it out simplifies your privacy position considerably.

    Can we standardize on an existing format?

    Adopting a community format helps interoperability and makes pooling with public data easier. Most teams extend one rather than adopting it unchanged, because the metadata layer is where their own requirements live.

    The log schema is the cheapest thing to get right at the start and the most expensive thing to retrofit. If you are designing one, or inherited one you do not trust, send us the shape of it and we will tell you what will hurt at scale.


    Pingal Mukherjee

    Pingal Mukherjee · Manager, Presales & Bid Management

    Pingal scopes data collection programs for humanoid, surgical, and industrial robotics clients, translating ML requirements into collection specifications for presales and bid teams, and writes about data strategy and sim-to-real transfer.

    Contact form

    Or just fill this out

    We’ll route your message to the right inbox and respond within one business day.