Frozen v8 gate-fix bundle (human_to_robot backbone, seed 42) — 2026-09-22

csv/
  hand_train.csv, hand_val.csv   — human hand pretraining data
                                    (Data_top20/ambulance_assembly_new/{train,val}.csv)
  aloha_train.csv, aloha_val.csv, aloha_test.csv
                                  — ALOHA robot finetune/eval splits (aloha_data/*.csv)
  aloha_cp.csv                   — conformal prediction cal/temp data (aloha_data/cpdata/data.csv)

weights/
  backbone/best_model_edit.pth   — MS-TCN++ backbone (14L, hand-pretrained + ALOHA-finetuned,
                                    human_to_robot, seed 42). label_mapping.json must sit
                                    beside it (load_model reads it from the ckpt's dirname).
  transformer_corrector/best_model.pt
                                  — gap-aware Transformer corrector (26-D features,
                                    human_to_robot backbone, seed 42)
  reliability/final_reliability_analyzer.pkl
                                  — reliability analyzer + transition matrix (flag threshold 0.39)

configs/
  base_label_mapping_22.json, class_index_names_23.json — canonical 22-class label maps

features/aloha/
  Pre-extracted X3D features (one .npy + one _labels.npy per video_id, 284 files, ~1.5G).
  This is what --feat should point at. Covers every video referenced by the CSVs above
  (hand + aloha train/val/test + cp) — NOT raw frame images (those are ~637G and not
  included; only needed to re-extract features from scratch or render camera footage).

code/
  Flat copy of the two directories the pipeline imports from:
    code/*.py     — top-level MSTCN++ scripts (model.py, etc.), 29 files
    code/v2/*.py  — v2 pipeline scripts (dump_demo_frame_csv.py, evaluate_transformer_final_test_26d.py,
                    tune_transformer_gate.py, reliability_analyzer.py, build_neural_dataset.py, ...), 165 files
  Scripts expect this same flat layout: code/ as MSTCN++ root, code/v2/ as the v2/ subdir,
  since v2 scripts do sys.path.insert(0, v2_dir) and sys.path.insert(0, dirname(v2_dir)).
  Does NOT include the many experiment-output subdirectories that normally live under v2/
  on disk (those are results, not code).

To reproduce the per-video CSV dump this was built for, from inside code/v2/:
  python dump_demo_frame_csv.py --seed 42 --videos <video_ids> \
      --ckpt ../weights/backbone/best_model_edit.pth \
      --feat ../features/aloha \
      --test-csv ../csv/aloha_test.csv \
      --reliability-pkl ../weights/reliability/final_reliability_analyzer.pkl \
      --transformer-ckpt ../weights/transformer_corrector/best_model.pt \
      --label-map ../configs/base_label_mapping_22.json \
      --class-names ../configs/class_index_names_23.json \
      --out out.csv --i-confirm-test
(Path args above are relative to code/v2/; adjust if you place things elsewhere.)

Frozen gate config used with these weights: mode E, tau=0.7, transition_margin=0.0
(human_to_robot backbone). See FROZEN_STATE_v8.md in the v2/ pipeline dir for full provenance.

This is seed 42 / human_to_robot only — the config dump_demo_frame_csv.py uses by default.
Seeds 123/456 and the robot_only backbone are not included; ask if you need those too.
