Market intelligence, with every claim sourced
Point-in-time events, filings, prices, macro releases, anomaly detection, and a timestamped evidence trail.
Read the design dossier →I build the full loop: data, models, evaluation, inference, and the product a person actually uses. This page keeps measured work separate from the systems I am designing next.
Every item here has a repository, a working artifact, measured results, or all three. Use the filters to follow one part of the stack.
A browser tool that adds persistent object tracks, masks, grasp points, and phase boundaries to LeRobot episodes. It turns a gripper-width trace into reach, transport, and retreat segments without a learned boundary model, then reserves SAM for the pixel work that actually needs one.
Matched 50-seed evaluation 02 · Robot policyA 262M-parameter manipulation policy trained on PushT and compared with the published checkpoint under identical rollout seeds, with confidence intervals and failure analysis.
Planning prototype · 40 paired seeds 03 · AutonomyExact signed-distance fields, sphere-traced collision checks, and controlled comparisons of RRT, RRT*, and Informed RRT*. The page includes a browser visualization and states what remains outside the closed loop.
End-to-end training run complete 04 · Model trainingLicense-gated data, full BF16 FSDP2 training on four H100s, a retention gate that caught a real HumanEval regression, FP8 serving, and GRPO with the compiler as the reward.
Clean and real-meeting benchmarks 05 · AudioWhisper plus an open diarization pipeline, WER and DER implemented from scratch, calibration on held-out data, and browser-tab transcription about five seconds behind live.
Recall / latency / memory sweep 06 · RetrievalEmbedding quality is measured independently from ANN fidelity, then five FAISS index families are compared. HNSW reaches 0.991 recall@10 at 0.23 ms in the reported run.
280 Tesseract evaluations · VLMs pending 07 · Document AI evaluationA provider-neutral benchmark with exact synthetic truth, 13 seeded corruption families, durable page records, and one honestly scoped local baseline. Real model runs and adapter contract tests stay visibly separate.
Engine verified · learning run in progress 08 · Reinforcement learningA legal move engine checked with perft, a browser game, and an AlphaZero-style self-play path. The case study separates the verified engine from the in-flight learning milestone.
Simulated portfolio with planted bias 09 · Responsible MLA synthetic 800,000-application portfolio where the redlining proxy is known in advance. The audit catches it and rejects the higher-Gini model, so fairness is tested against ground truth.
The recurring pattern is to preserve provenance from raw data to user-visible decision and attach an evaluation gate to every handoff.
Episodes, audio, documents, filings, labels, licenses, and timestamps.
Fine-tuning, imitation, RL, retrieval, and the simplest credible baseline.
Matched seeds, held-out environments, calibration, regressions, and abstention.
Browser interfaces, streaming inference, latency budgets, observability, and cost.
These are implementation dossiers, not completed-project claims. Each page defines the artifact, data contract, evaluation, safety boundary, and the evidence required before the status changes.
Point-in-time events, filings, prices, macro releases, anomaly detection, and a timestamped evidence trail.
Read the design dossier →Food, labs, vitamins, exercise, sleep, family history, and ancestry-aware evidence without turning a wellness app into an unvalidated diagnosis engine.
Read the design dossier →Fine-tuning, synthetic data, serving, and quality-versus-cost gates across dated, pinned model revisions.
Read the model playbook →A story-first shot graph joining ComfyUI generation with traditional editing, compositing, color, sound, and continuity review.
Read the production plan →Codex and Claude as planner, implementer, reviewer, and critic behind one traceable task protocol and test gate.
Read the harness design →LeetGPU kernels organized by the architecture idea they prove, from coalescing through fused LLM blocks and simulation.
Read the kernel roadmap →The long-form library is reference material behind the projects, not a substitute for them.
Behavior cloning, DAgger, diffusion policies, VLA models, world models, sim-to-real, and evaluation.
Data mixtures, fleet collection, interventions, held-out environments, and the deployment flywheel.
HBM, tensor cores, systolic dataflows, precision, sparsity, and measured H100 rooflines.
A dated, source-linked timeline of the ideas that converged across language, vision, audio, robotics, RL, and compute.
I am most useful where a model has to become a dependable system: robotics and embodied AI, model training and serving, reinforcement learning, multimodal products, or the accelerated compute underneath them.