Meeting notes — 2026-09-30
Updated: 2026-09-30
Recorded discussion, 16:23. Topic: the water-content throwing experiment and what the unifying research contribution should be. Transcribed from audio; several technical terms were garbled by speech recognition and are corrected inline.
Prior session: 2026-09-23.
#Water-bottle flipping experiment
The plan: vary water content in the bottle so the required throw differs, then test whether a policy picks up on it. Pressed on purpose — the point is whether tactile lets the policy infer mass/fill level and adapt, and whether a vision-only policy can do the same.
Consensus: vision has a plausible path, since the robot lifts the bottle before throwing so fill level is partly observable, and a transparent bottle makes it easier — but it is harder than tactile, and mass estimation from vision is acknowledged prior work.
Scope clarification: this is not in-context learning. It only tests whether tactile observes the variation and the policy shifts behaviour accordingly. That is the easier, worthwhile first result.
Decision: worth running; a transparent bottle is acceptable.
#Frisbee throwing — dropped
Raised as an alternative task, rejected on hardware grounds: it needs wrist rotation the current rig cannot produce.
#The unifying story
The tasks are appealing but the technical contribution is unclear. Framing offered: dynamic dexterous manipulation is an open niche — "nobody's really doing dexterous and dynamic". Two possible shapes:
- Attempt hard tasks, hit concrete obstacles, and the fix becomes the contribution.
- Find a unifying technical approach up front.
#Idea A — time-warping human demonstrations
When a human demonstrates something fast you cannot absorb it all at once; you would slow down the critical parts and replay them. Can the equivalent be done for policy learning, non-RL — artificially slow a demonstration, learn from it, then progressively speed up?
- Naive versions were pre-emptively dismissed: dropping alternate actions in a chunk, or upsampling action resolution near contact-rich segments. Action-chunk timing and control frequency are too tightly coupled to data fit.
- Pushback: dynamic manipulation is by definition not quasi-static, so slowing it changes the outcome.
- Counter-suggestion: run the speed curriculum in simulation — train at ~300 Hz decision-making, then gradually raise simulation speed.
- Related prior art: a speech-domain loss (said "CPC", describing CTC) that ignores silences and filler and scores only meaningful tokens. Analogous in spirit to weighting only the important segments of a trajectory.
#Idea B — solve for the action manifold from physics
To move an object along a target trajectory, a determinate amount of work must be delivered in a given time, and that work can only come from the robot hand. Therefore:
- Infinitely many action sequences produce the same object trajectory; a demonstration provides one solution to that constrained problem.
- Could the manifold of equivalent solutions be learned, and interpolated to find others?
- Practical route: capture human video, extract object state/trajectory plus human pose. Important distinction — plain BC on captured object+pose trajectories is a different problem from solving for the actions that cause the object motion.
- Useful framing: a poor human demonstrator can still show the correct poses even when the velocities are wrong. Capture the pose sequence, then inject a correct velocity profile afterwards and learn from that.
#Proposed north star
The extreme form of dexterous manipulation: specify exactly what object pose is required at each time step and require the robot to hit it. Currently unsolved — sim2real would not crack it, and arguably no human can do it either. Bottle flipping is a strictly easier instance, so it is a first rung on that ladder.
#Decisions
- Try the sim2real route first for bottle flipping, to establish feasibility — it will produce insights or at least more to discuss.
- Agreement the task is worth solving regardless of eventual framing.
- Implementation details to be worked out offline.
- Progress over the preceding two weeks called out as very impressive.
#Post-meeting notes
A smaller exchange after the formal close:
- Scepticism that one sub-problem needs a learning solution at all — it can be hand-defined, as in another member's project (semantic links defined per region). Easy to set up and not a bottleneck.
- Recommendation to look at real2sim first, reusing an existing foundation built for a teammate's project, which reportedly tracks thrown objects like this well.
- Continued interest in trying tactile BC, noting neither party had invested much in it yet; agreement to try since it is cheap.
- Composing skills and solving in MuJoCo once object state capture is in place.