Parallel simulated tasks
Three manipulation tasks across randomized objects, textures, lighting, cameras, and physics. Record every seed and environment revision.
I am building the full embodied-learning loop: episode-centric robot data, imitation and VLA policies, simulation and world models, motion planning, and evaluation in environments the policy did not see.
Each status describes the artifact that exists today, not the final system I eventually want to build.
Tracklet adds persistent object identity, masks, grasp points, and subtask phases to LeRobot demonstrations. Its key shortcut is physical: gripper width reveals grasp and release boundaries, so deterministic signal processing proposes the episode structure and SAM handles only the pixel segmentation.
A 262M-parameter PushT policy trained from scratch and evaluated beside the published reference under identical rollout seeds and confidence intervals.
Planning prototype · 40 paired seeds Motion planningExact signed-distance fields, sphere-traced collision checks, and controlled RRT comparisons in a known map. The estimator and closed-loop flight controller remain unbuilt.
The next experiment should reuse the same episodes, simulators, metrics, and deployment path instead of becoming another isolated demo.
Observations, actions, language, proprioception, contact, and provenance.
Object tracks, spatial tokens, latent dynamics, action chunks, and uncertainty.
Behavior cloning first, then model-predictive planning and policy improvement.
New objects, layouts, viewpoints, perturbations, latency, safety, and sim-to-real gap.
A serious next flagship needs a running simulator, artifact-backed evaluations, and a narrow real-world transfer target.
Three manipulation tasks across randomized objects, textures, lighting, cameras, and physics. Record every seed and environment revision.
Compare a compact action-chunking or diffusion baseline with a pretrained VLA adapter under the same episodes and action representation.
Predict future latent observations and rewards, show predicted-versus-actual rollout strips, then use the model for receding-horizon selection.
Success heatmaps on held-out objects and layouts, calibration of predicted success, compute/success Pareto curves, and failure videos.
Long-form notes are supporting depth, not evidence that an unlinked product already exists.
USD, PhysX, RTX sensors, hardware constraints, installation, and the cases where a purpose-built simulator is the better tool.
GPU-parallel robot learning, environment APIs, version compatibility, training, and the Isaac Gym-to-Lab lineage.
Behavior cloning, DAgger, diffusion policies, VLA models, world models, sim-to-real, and evaluation.
Fleet data, interventions, mixture design, held-out worlds, and the improvement flywheel.
Latency, control interfaces, observation/action recording, safety, and demonstration quality.
Belief state, planning, risk, and the boundary between learned and model-based control.
USD stage, the SimulationApp bootstrap, sensors, and the RT-core requirement that decides which machines can run it at all.
GPU-parallel robot learning, the 2.0 import rename that breaks most tutorials online, and why a datacenter GPU cannot run it.