World models, VLA, and simulation
A robot learning loop from demonstrations through learned dynamics, planning, and held-out-world evaluation.
See the embodied-AI roadmap →My work connects robot data and simulation, multimodal models, reinforcement learning, inference systems, and the product interface. I care about controlled experiments, visible failure modes, and getting the whole system into a person's hands.
Shipped interfaces and measured experiments are labeled separately from active builds.
A browser-first annotation workspace for persistent tracks, masks, grasp points, and manipulation phases. It uses gripper proprioception for deterministic phase boundaries and a compact browser SAM model only where pixels require it.
Measured
Robot policy
A 262M-parameter manipulation policy compared against the published reference under the same PushT rollout seeds.
Planning prototype AutonomyExact distance fields, sphere-traced collision checks, and paired RRT-family comparisons. State estimation and closed-loop flight remain future work.
Training run complete Open-model systemsLicense-gated data, full FSDP2 fine-tuning, a retention regression that stopped the launch gate, FP8 serving, and GRPO.
Measured Audio intelligenceWhisper plus open diarization, with WER and DER written from scratch and evaluated on both clean speech and real meetings.
Live at gradientsmith.ai Model evaluationA verifier-first evaluation and post-training platform. Deterministic sandboxed verifiers instead of an LLM judge, an adversary that mines its own counterexamples, cost-aware routing across models, and RS-SFT plus GRPO with a reward-hacking monitor.
Useful ML work is a connected chain; an accuracy number without data provenance or deployment constraints is incomplete.
Episodes, demonstrations, audio, filings, health records, licenses, and timestamps.
VLA policies, world models, retrieval, fine-tuning, self-play, and credible baselines.
Matched seeds, held-out worlds, point-in-time splits, calibration, safety, and abstention.
Kernels, serving, cost, observability, and an interface that makes uncertainty legible.
These are transparent implementation dossiers, not claims of completed products.
A robot learning loop from demonstrations through learned dynamics, planning, and held-out-world evaluation.
See the embodied-AI roadmap →Dated choices across Kimi K3, Qwen3.6/3.7, and GLM-5.2, plus fine-tuning, data, serving, and quality-per-dollar gates.
Read the playbook →Codex and Claude behind one task protocol with isolated work, trace replay, test gates, routing, and cost accounting.
Read the harness design →LeetGPU work organized by memory behavior and fused LLM blocks, then connected to a quantized inference accelerator.
Open the roadmap →SEC filings, macro releases, prices, event graphs, regime shifts, and leakage-safe temporal evaluation.
Read the system design →Meals, labs, vitamins, activity, family history, and genetics, designed as transparent wellness support, not diagnosis.
Read the safety-first design →I studied electrical engineering at UT Austin and computer science at Stanford. I am most useful where models, systems, and product decisions meet: embodied AI, multimodal learning, reinforcement learning, model infrastructure, or accelerated computing.