Back to Shotwell

Case study

Shotwell x Ultra: Turning Robot Teleoperation Data Into Reliable Training Data

Introduction

Ultra is one of the leading robotics companies deploying robots in factories and logistics centers across the US. In these settings, teleoperators produce large volumes of data that are used for policy training. Across thousands of hours, teleoperators make mistakes: they drop items, pack things incorrectly, or place things in the wrong containers.

Shotwell partnered with Ultra to label and structure teleoperation episodes from real robot camera feeds, turning raw video into searchable, high-signal training data. With Shotwell, Ultra can identify clean demonstrations and build high-quality datasets for policy training.

The Challenge

Ultra's robot policies are trained on teleoperation data collected in real, unscripted environments. This data is valuable because it captures the messiness of production that scripted demos lack. But raw teleoperation video is difficult to use at scale.

Before Shotwell, Ultra's team had to spend significant engineering time reviewing episodes, giving feedback to operators, and manually identifying which demonstrations were suitable for training. One engineer could keep up with feedback and quality management for the output of 10 teleoperators. As the teleoperation operation grew, this workflow became a bottleneck.

Ultra needed a way to answer questions like:

Without dense labels, these questions required manual review. With Shotwell, they became database queries.

The Solution

Shotwell provided Ultra with dense event and subtask annotations across its teleoperation data. Instead of treating each episode as a single opaque video, Shotwell decomposed episodes into meaningful subtasks with attributes and failure mode detection per subtask.

Input Ultra Teleoperation Episode

Output Shotwell Annotations

    For Ultra's tower stacking toy task, Shotwell annotated each "Pick", "Set", and "Unstack" subtask the robot performed. For each subtask, Shotwell provided color attributes and failure attributes.

    For Ultra's other tasks, there is a separate vocabulary of subtasks and a separate list of attributes per subtask. Ultra can then slice training data along the dimensions they care about to curate the right training set. Noisy or failed attempts can be excluded or filtered in different ways for different policy experiments. Subtask boundaries can be used to train more precise policies. Operator behavior can be measured consistently over time.

    The result is a data layer that makes robotics training data more transparent, searchable, and useful.

    The Process

    Shotwell and Ultra built the annotation workflow in three stages.

    1. SOP Creation

    The first step was to define what "good" looked like.

    Shotwell worked with Ultra to translate task goals and substeps into a concrete annotation SOP. This included the task schema, event definitions, and success criteria.

    2. Manual Annotation and Schema Alignment

    Shotwell began with manual annotation to align closely with Ultra's internal understanding of each task.

    This stage resolved ambiguity such as what motion counts as a retry, when a subtask begins, and what exactly is considered a failure. Going through many examples of episodes and annotating them manually allowed Shotwell to quickly align with Ultra on these questions.

    3. Model Training + Human-in-the-Loop

    Once the schema was stable, Shotwell trained models to identify events and segment boundaries on an initial batch of approved human annotations.

    Shotwell then moved the workflow into a steady-state model-plus-human review process. In tasks where the models had lower confidence, the models provided pre-annotations that a human annotator could quickly fix.

    As more episodes were annotated, the model was further trained and developed higher confidence across the subtasks. At this stage, humans were still used to check for data drift and perform model QA. Humans could also be flagged to re-annotate low confidence segments of otherwise high confidence episodes. This annotation flywheel continuously improves model performance and speeds up the overall annotation process.

    Shotwell's approach scaled quickly to Ultra's load to provide 24 hour turnaround times, which were particularly useful for operator feedback.

    Outcomes

    With Shotwell, Ultra transformed teleoperation data review from a manual bottleneck into a scalable data operation.

    One engineer who previously struggled to manage feedback and quality review for around 10 operators can now support 60+ operators using Shotwell's annotation and review pipeline. That unlocks a major scaling advantage: Ultra can grow teleoperation data collection without growing engineering overhead at the same rate.

    With higher quality annotations, Ultra was able to improve their policy success rate and task adherence.

    Shotwell also delivered materially faster turnaround than competing annotation solutions. Ultra had tried two approaches prior to Shotwell. They had teleoperators perform task annotation as they performed actions, but those annotations were lower quality, around 80% accurate, and far less granular than what Shotwell provided. They also tried hiring labeling companies that were more expensive, lower fidelity in annotation, and unable to provide single day turnaround times because they relied entirely on offshore manual labor.

    When it comes to training robot policies, data quality is everything, and Shotwell guarantees it.