AI / ML systems engineer

I build machines that learn, decide, and move.

My work connects robot data and simulation, multimodal models, reinforcement learning, inference systems, and the product interface. I care about controlled experiments, visible failure modes, and getting the whole system into a person's hands.

4 × H100 Full BF16 fine-tune, FP8 deployment, and compiler-rewarded RL.
99.1% recall@10 At 0.23 ms p50 for the selected HNSW operating point.
50 matched seeds Robot-policy comparisons with confidence intervals.
40 paired seeds Controlled RRT-family planning benchmarks.
Featured work

Robotics first, evidence throughout

Shipped interfaces and measured experiments are labeled separately from active builds.

Working local demo · packaging now Robot data

Tracklet turns robot episodes into training-ready supervision

A browser-first annotation workspace for persistent tracks, masks, grasp points, and manipulation phases. It uses gripper proprioception for deterministic phase boundaries and a compact browser SAM model only where pixels require it.

Tracklet annotation interface showing a robot episode, timeline, object tracks, masks, and grasp controls
LeRobotSAM 2.1 ONNX Runtime WebNext.js
Measured Robot policy

Diffusion policy from scratch

A 262M-parameter manipulation policy compared against the published reference under the same PushT rollout seeds.

Diffusion policyLeRobotPyTorch
Planning prototype Autonomy

GPS-denied motion planning

Exact distance fields, sphere-traced collision checks, and paired RRT-family comparisons. State estimation and closed-loop flight remain future work.

C++17CUDARRT*
Training run complete Open-model systems

An 8B Rust coder, end to end

License-gated data, full FSDP2 fine-tuning, a retention regression that stopped the launch gate, FP8 serving, and GRPO.

Qwen3-8BFSDP2vLLM
Measured Audio intelligence

Who said what, seconds behind live

Whisper plus open diarization, with WER and DER written from scratch and evaluated on both clean speech and real meetings.

WhisperDiarizationStreaming
Live at gradientsmith.ai Model evaluation

gradientsmith scores models on tests they cannot argue with

A verifier-first evaluation and post-training platform. Deterministic sandboxed verifiers instead of an LLM judge, an adversary that mines its own counterexamples, cost-aware routing across models, and RS-SFT plus GRPO with a reward-hacking monitor.

Verifier-firstGRPOQwen2.5-CoderNext.js
Systems view

The loop I like to own

Useful ML work is a connected chain; an accuracy number without data provenance or deployment constraints is incomplete.

01 · Data

Observe and label

Episodes, demonstrations, audio, filings, health records, licenses, and timestamps.

02 · Model

Learn and simulate

VLA policies, world models, retrieval, fine-tuning, self-play, and credible baselines.

03 · Evaluate

Test the decision

Matched seeds, held-out worlds, point-in-time splits, calibration, safety, and abstention.

04 · Operate

Deliver efficiently

Kernels, serving, cost, observability, and an interface that makes uncertainty legible.

Active portfolio roadmap

What I am building toward

These are transparent implementation dossiers, not claims of completed products.

Build dossier

World models, VLA, and simulation

A robot learning loop from demonstrations through learned dynamics, planning, and held-out-world evaluation.

See the embodied-AI roadmap →
Build dossier

Frontier open-model cost lab

Dated choices across Kimi K3, Qwen3.6/3.7, and GLM-5.2, plus fine-tuning, data, serving, and quality-per-dollar gates.

Read the playbook →
Critical next flagship

My own AI coding harness

Codex and Claude behind one task protocol with isolated work, trace replay, test gates, routing, and cost accounting.

Read the harness design →
Portfolio roadmap

GPU kernels to accelerator architecture

LeetGPU work organized by memory behavior and fused LLM blocks, then connected to a quantized inference accelerator.

Open the roadmap →
Build dossier

Market intelligence with provenance

SEC filings, macro releases, prices, event graphs, regime shifts, and leakage-safe temporal evaluation.

Read the system design →
Build dossier

Personal health intelligence

Meals, labs, vitamins, activity, family history, and genetics, designed as transparent wellness support, not diagnosis.

Read the safety-first design →
Open to AI / ML / robotics systems roles

Looking for an engineer who crosses the stack?

I studied electrical engineering at UT Austin and computer science at Stanford. I am most useful where models, systems, and product decisions meet: embodied AI, multimodal learning, reinforcement learning, model infrastructure, or accelerated computing.