Uncertainty-Aware Human-to-Robot Intention Prediction for Assisted Teleoperation

Abstract

In assisted teleoperation for human-robot collaboration, accurate intention prediction is critical for enabling timely and reliable robotic assistance during long-horizon manipulation and assembly tasks. Robot teleoperation demonstrations are costly and hardware-limited, whereas human demonstrations are easier to collect and provide rich temporal structure. To address this challenge, we propose an uncertainty-aware human-to-robot intention prediction framework that combines: (1) hierarchical transfer learning, where MS-TCN++ is pretrained on human hand demonstrations and fine-tuned on limited robot teleoperation data; (2) a conformal prediction module that provides frame-level prediction sets with statistical coverage guarantees for reliable uncertainty quantification; and (3) reliability-gated segment correction, where a reliability analyzer flags uncertain segments, a gap-aware Transformer proposes corrected labels, and a gate accepts a change only when it is confident and transition-consistent.

Experiments on robot assembly demonstrations with 22 action classes show that human-to-robot fine-tuning improves the robot test-set Edit score from 55.23 to 72.51 (mean over 3 seeds) using only 9 robot demonstrations — a gain of 17.28 points over training from scratch. Gated correction further raises Edit to 73.45 without lowering frame accuracy, and conformal prediction sets meet their target coverage. These results show that human demonstrations provide scalable pretraining data for robust, uncertainty-aware robot action segmentation.

Overview

Framework overview

Our framework: (1) MS-TCN++ pretrained on human demonstrations is fine-tuned on robot data; (2) conformal prediction quantifies per-frame uncertainty; (3) uncertain segments are re-labelled by a gap-aware Transformer corrector with gated acceptance.

Method

Three components work together to produce reliable, uncertainty-calibrated action segmentation.

1

Cross-Domain Transfer

MS-TCN++ is pretrained on 83 UMI hand videos encoding action transition priors, then fine-tuned on 18 ALOHA robot videos (9 demonstrations). Both domains share a 22-class vocabulary and 192-dim X3D-M features, enabling complete weight transfer without architectural changes.

2

Conformal Prediction

Temperature-scaled softmax probabilities are calibrated on a held-out split to produce prediction sets with finite-sample marginal coverage guarantees. Two variants are evaluated: standard (marginal) and class-conditional conformal prediction with per-class thresholds.

3

Gated Segment Correction

A reliability analyzer flags segments likely to be wrong. A gap-aware Transformer, pretrained on human data and fine-tuned on robot data, proposes a new label for each segment. A change is accepted only if the segment was flagged, the proposal confidence is ≥ 0.7, and the new label is transition-compatible with its neighbours.

Why human demonstrations?

The 22×22 action transition matrix has 484 entries. The 18 robot training videos contain only 436 observed transitions — insufficient to learn long-range assembly ordering. The 83 human training videos contain 4,025 transitions, and pretraining on them transfers these priors via fine-tuning.

Human and Robot Datasets

Hand demo

Source Domain — UMI

83 training + 20 validation hand videos of toy car assembly captured with an egocentric GoPro camera on a handheld UMI gripper.

ALOHA demo

Target Domain — ALOHA

Teleoperated demonstrations recorded with dual wrist cameras (480×640). Split: 18 train (9 demos) / 26 val (13 demos) / 34 test (17 demos) / 46 calibration videos.

Results

Evaluated on the robot test set (34 videos, 17 demonstrations) using frame accuracy, Edit score, and F1 at multiple overlap thresholds. Mean ± std over 3 seeds.

73.45
Edit Score
Human→Robot + gated corrector
+17.3
Edit Score Gain
over robot-only baseline
34.2%
Frame Accuracy
after gated correction
9
Robot Demos
18 training videos

Transfer Learning

Model Edit ↑ Acc % ↑ F1 @10 ↑ F1 @25 ↑ F1 @50 ↑
Robot-Only (trained from scratch) 55.23 ± 5.9628.48 ± 1.2632.2324.0110.97
Robot-Only + Gated Corrector 55.99 ± 4.2528.82 ± 0.9432.9924.6911.25
Human → Robot (ours) 72.51 ± 2.1934.11 ± 0.5538.5230.8116.86
Human → Robot + Gated Corrector (ours, best) 73.45 ± 2.81 34.19 ± 0.60 38.58 30.83 16.89
Key Takeaway — Edit Score
55.23 Robot-Only
→
72.51 Human → Robot
+17.28

Conformal Prediction

✓
Both CP variants meet the 93% and 97% coverage targets in every setting
17–27%
Smaller prediction sets with class-conditional CP vs. standard CP
~21
Standard CP mean prediction set size out of 22 classes (Human→Robot)
Standard CP results

Standard CP

Class-conditional CP results

Class-Conditional CP

CP Method Target Level Robot-Only Coverage Human→Robot Coverage + Gated Corrector Coverage Human→Robot Set Size
Standard CP (marginal) 93%96.3%95.8%96.5%21.0 / 22
97%99.3%97.9%98.4%21.5 / 22
Class-Conditional CP 93%94.9%95.3%95.4%17.2 / 22
97%97.4%97.1%97.3%17.9 / 22

Segmentation Results

Segment-level predictions for teleoperated robot video

Human→Robot + gated corrector predictions on test videos v220 / v221

StartEndDuration Ground Truth Action Predicted Action IoUResult
415876461pick up screwposition screw on first wheel58.6%incorrect
8761311435position screw on first wheelposition screw on first wheel32.1%correct
13111430119pick up first wheelpick up first wheel36.0%correct
143025741144position first wheelposition first wheel74.8%correct
25743172598pick up electric screwdriverpick up electric screwdriver33.4%correct
31723996824position screwdriver bitscrew first wheel with screwdriver44.8%incorrect
39964194198screw first wheel with screwdriverpick up screw19.7%incorrect
41944383189put down electric screwdriverpick up screw47.6%incorrect
43834604221pick up screwpick up screw21.7%correct
46045075471position screw on second wheelposition screw on second wheel44.1%correct
50755223148pick up second wheelposition screw on second wheel13.9%incorrect
52236052829position second wheelposition second wheel26.0%correct
60526618566pick up electric screwdriverposition second wheel56.8%incorrect
66186834216position screwdriver bitpick up electric screwdriver41.4%incorrect
68347680846screw second wheel with screwdriverscrew second wheel with screwdriver42.8%correct
76807972292put down electric screwdriverpick up screw64.0%incorrect
79728244272pick up screwposition screw on side roof30.6%incorrect
82448982738position screw on side roofposition screw on side roof54.3%correct
89829239257pick up electric screwdriverposition screwdriver bit45.9%incorrect
92399716477position screwdriver bitscrew side roof with screwdriver24.8%incorrect
971610563847screw side roof with screwdriverscrew side roof with screwdriver71.4%correct
1056310719156put down electric screwdriverput down electric screwdriver61.3%correct

11 correct / 22 labelled segments  |  Human→Robot + gated corrector (seed 42)  |  Video v220  |  Predicted action = best-overlapping predicted segment