Selected work

Embodied AI and efficient learning systems.

I build the full loop: data, models, evaluation, inference, and the product a person actually uses. This page keeps measured work separate from the systems I am designing next.

4 × H100 Full BF16 fine-tune, FP8 export, served inference, and compiler-rewarded RL.
99.1% at 0.23 ms ANN fidelity and single-query latency for the selected HNSW operating point.
Matched seeds Robot policies and planners compared under paired evaluation, not cherry-picked runs.
C++ / CUDA → UI Systems work that runs from kernels and control loops through browser products.
Evidence first

Case studies

Every item here has a repository, a working artifact, measured results, or all three. Use the filters to follow one part of the stack.

Working local demo · packaging now 01 · Robot data

Tracklet: episode-centric annotation for robot learning

A browser tool that adds persistent object tracks, masks, grasp points, and phase boundaries to LeRobot episodes. It turns a gripper-width trace into reach, transport, and retreat segments without a learned boundary model, then reserves SAM for the pixel work that actually needs one.

LeRobot ONNX Runtime Web Robot proprioception Next.js
Matched 50-seed evaluation 02 · Robot policy

Diffusion policy from scratch

A 262M-parameter manipulation policy trained on PushT and compared with the published checkpoint under identical rollout seeds, with confidence intervals and failure analysis.

Diffusion policy LeRobot PyTorch
Planning prototype · 40 paired seeds 03 · Autonomy

GPS-denied motion planning

Exact signed-distance fields, sphere-traced collision checks, and controlled comparisons of RRT, RRT*, and Informed RRT*. The page includes a browser visualization and states what remains outside the closed loop.

C++17 CUDA RRT*
End-to-end training run complete 04 · Model training

An 8B Rust coder, fully fine-tuned

License-gated data, full BF16 FSDP2 training on four H100s, a retention gate that caught a real HumanEval regression, FP8 serving, and GRPO with the compiler as the reward.

Qwen3-8B FSDP2 vLLM
Clean and real-meeting benchmarks 05 · Audio

Speaker-attributed transcription

Whisper plus an open diarization pipeline, WER and DER implemented from scratch, calibration on held-out data, and browser-tab transcription about five seconds behind live.

Whisper Diarization CTranslate2
Recall / latency / memory sweep 06 · Retrieval

Vector search as two separate experiments

Embedding quality is measured independently from ANN fidelity, then five FAISS index families are compared. HNSW reaches 0.991 recall@10 at 0.23 ms in the reported run.

FAISS BEIR Ray Data
280 Tesseract evaluations · VLMs pending 07 · Document AI evaluation

Build the OCR arena before the leaderboard

A provider-neutral benchmark with exact synthetic truth, 13 seeded corruption families, durable page records, and one honestly scoped local baseline. Real model runs and adapter contract tests stay visibly separate.

OCR VLMs Modal Robustness
Engine verified · learning run in progress 08 · Reinforcement learning

Chess that proves itself before it learns

A legal move engine checked with perft, a browser game, and an AlphaZero-style self-play path. The case study separates the verified engine from the in-flight learning milestone.

MCTS Self-play JavaScript
Simulated portfolio with planted bias 09 · Responsible ML

Credit risk with a provable audit

A synthetic 800,000-application portfolio where the redlining proxy is known in advance. The audit catches it and rejects the higher-Gini model, so fairness is tested against ground truth.

LightGBM SHAP Fairness
How I work

One loop, not disconnected demos

The recurring pattern is to preserve provenance from raw data to user-visible decision and attach an evaluation gate to every handoff.

01 · Observe

Data with provenance

Episodes, audio, documents, filings, labels, licenses, and timestamps.

02 · Learn

Models with baselines

Fine-tuning, imitation, RL, retrieval, and the simplest credible baseline.

03 · Decide

Gates and uncertainty

Matched seeds, held-out environments, calibration, regressions, and abstention.

04 · Deliver

Products and operations

Browser interfaces, streaming inference, latency budgets, observability, and cost.

Explicitly not shipped yet

Building next

These are implementation dossiers, not completed-project claims. Each page defines the artifact, data contract, evaluation, safety boundary, and the evidence required before the status changes.

Build dossier

Market intelligence, with every claim sourced

Point-in-time events, filings, prices, macro releases, anomaly detection, and a timestamped evidence trail.

Read the design dossier →
Build dossier

Personal health intelligence

Food, labs, vitamins, exercise, sleep, family history, and ancestry-aware evidence without turning a wellness app into an unvalidated diagnosis engine.

Read the design dossier →
Engineering playbook

Open-weight model cost lab

Fine-tuning, synthetic data, serving, and quality-versus-cost gates across dated, pinned model revisions.

Read the model playbook →
Build dossier

Controllable AI film pipeline

A story-first shot graph joining ComfyUI generation with traditional editing, compositing, color, sound, and continuity review.

Read the production plan →
Critical next flagship

A provider-neutral coding-agent harness

Codex and Claude as planner, implementer, reviewer, and critic behind one traceable task protocol and test gate.

Read the harness design →
Portfolio roadmap

GPU Kernel Lab and accelerator path

LeetGPU kernels organized by the architecture idea they prove, from coalescing through fused LLM blocks and simulation.

Read the kernel roadmap →
Open to AI / ML / robotics systems roles

Looking for someone who can cross the stack?

I am most useful where a model has to become a dependable system: robotics and embodied AI, model training and serving, reinforcement learning, multimodal products, or the accelerated compute underneath them.