Consent and Privacy in Egocentric Capture: A Practical Checklist

Consent and Privacy in Egocentric Capture

A camera on someone’s head records everything in front of them, including people who never agreed to be recorded. Egocentric data collection consent is therefore a design problem rather than a form to sign, and it is far cheaper to build in than to retrofit.

This guide covers the layers of consent egocentric capture requires, a ten-point practical checklist, why withdrawal has to be executable rather than promised, and what enterprise buyers will ask to see.

Table of contents

    A practical checklist rather than legal advice. Requirements differ by jurisdiction and by what you capture, so involve counsel before you start.

    The problem egocentric capture creates

    A fixed camera in a collection cell records a controlled space. Everyone in frame is there to be there.

    A head-mounted camera walking through a house records whoever is in the house. It records the neighbour at the door, the letter on the counter, the child in the next room, and the conversation happening nearby. The contributor consented to all of that. Nobody else did.

    This is not an edge case to be handled later. It is the normal operating condition of wearable capture, and it has to be designed for.

    The layers of consent

    Layer Who What it must cover
    Contributor The person wearing the rig What is captured, retention, who sees it, withdrawal
    Household Other adults in the space Obtained before capture, not after
    Bystander Visitors, passers-by Usually managed by redaction and exclusion zones
    Site Employer or property owner Workplace capture runs through existing consultation
    Subject matter Documents, screens, correspondence Incidental capture of confidential material

    The last row is the one most often missed. A rig pointed at a kitchen counter records the post. A rig in an office records screens. Neither is a person and both can be a disclosure.

    The checklist

    1. Written contributor consent in plain language, specific about retention period and about who may view footage.
    2. Household consent captured before the first session, with a record of who was covered.
    3. An easy, obvious pause control, and a culture where using it carries no penalty. If contributors feel pausing costs them work, they will not pause.
    4. Defined no-capture zones, agreed in advance rather than judged in the moment.
    5. Exclusion of environments with children unless you have a specific, reviewed protocol for it. Most programmes exclude rather than manage.
    6. Automated redaction at ingest for faces and identifiers, verified by sampling rather than assumed.
    7. Audio handled separately. It is frequently the most sensitive stream and the easiest to drop if you do not need it.
    8. Episode-level provenance, so a withdrawal request can actually be executed.
    9. Defined retention for the raw layer, enforced automatically rather than by intention.
    10. Access control by role and by region, because viewing is processing, per data residency.

    Withdrawal has to be executable

    Nearly every consent form promises that a contributor can withdraw. Very few pipelines can actually deliver it.

    Once footage has been through annotation, pooled into training splits, and used to produce model checkpoints, removing one contributor’s data is only possible if provenance was tracked at episode level from the beginning. Retrofitting that into a corpus of thousands of sessions is close to impossible.

    • Tag every episode with contributor and session identifiers that persist through every derived artifact.
    • Record which training splits an episode entered, so the blast radius of a withdrawal is knowable.
    • Decide the model position in advance. Whether withdrawal requires retraining is a commercial and legal question, and it should be answered in the consent language rather than improvised later.
    • Make deletion evidenced, not asserted.

    The honest version of this is to be specific in the consent form about what withdrawal does and does not reach, rather than implying more than the pipeline can deliver.

    Workplace capture is a different process

    Recording people at work does not run through your consent form. It runs through the employer’s existing employee consultation process, and in many jurisdictions through works councils or union agreements.

    Practical consequences: the site owner, not you, leads the conversation; lead times are longer than technical readiness suggests; and redaction requirements are often stricter because the recordings can bear on individual performance. This is covered further in industrial teleoperation data.

    For public-space capture, our aerial perception case study covers a programme run under those constraints, reporting a 3x detection rate from paired drone-and-ground capture.

    What buyers will ask to see

    • The consent template itself, not a description of it
    • Evidence that redaction runs automatically and is verified
    • Retention schedules and how deletion is proven
    • Who can access raw footage and from where
    • Whether children or clinical settings were ever in scope
    • Whether subcontractors touched the data

    Our own position is set out on the operator consent page, and the wider procurement picture in what enterprise procurement actually asks.

    Frequently asked questions

    Is blurring faces enough?

    It reduces risk and rarely achieves anonymisation on its own. Voice, gait, home interiors, and visible documents can all identify. Treat redaction as one control among several.

    Do we need consent from someone who walks past a delivery robot?

    Individual consent is usually impractical in public space, and obligations still apply. Programmes generally rely on redaction, minimisation, retention limits, and clear signage where required, guided by local rules.

    Can contributors record in shared housing?

    Only with consent from the other adults, obtained in advance. If that cannot be reliably managed, most programmes exclude those environments.

    How long should raw footage be kept?

    As briefly as the work allows. Derived data supports most training, so the raw layer can often have a much shorter life than the dataset itself.

    Consent designed in at the start costs very little. Consent retrofitted into a live corpus is expensive and sometimes impossible. If you want a capture protocol reviewed before collection begins, tell us what you are building.


    manish

    manish ·

    Contact form

    Or just fill this out

    We’ll route your message to the right inbox and respond within one business day.