AI & model systems

Models are useful when the surrounding system is measurable.

I work across data curation, full and parameter-efficient fine-tuning, reinforcement learning, retrieval, multimodal inference, evaluation, and the decision interface around the model.

8.19B parametersFull BF16 checkpoint trained with FSDP2 across four H100s.
23% exact matchHeld-out Rust generation after SFT; retention regression also reported.
0.38% clean DERCalibrated clean-speech result; real-meeting reversal kept visible.
0.991 recall@10HNSW fidelity at the selected 0.23 ms p50 operating point.
Evidence first

Measured case studies

Results include the regressions, domain shifts, and small-sample limits that change what the headline means.

End-to-end run complete Training + RL + serving

Full fine-tuning an 8B Rust coder, then paying the retention bill

A license-gated 14,080-pair corpus, full BF16 FSDP2 training on four H100s, FP8 export and vLLM serving, then GRPO with the compiler as verifier. Domain exact match reached 23%, while a 7.3-point HumanEval regression tripped the retention gate.

Qwen3-8BFSDP2GRPOvLLM
Clean + real-meeting evaluation Audio

Speaker-attributed transcription

WER and DER implemented from scratch, held-out calibration, a browser-tab streaming path, and the honest result that the clean-benchmark ranking reverses on AMI meetings.

WhisperCTranslate2Diarization
Six embedders · five indexes Retrieval

Separate embedding quality from ANN fidelity

Embedding models are judged against relevance; FAISS indexes are judged against exact neighbors. HNSW reaches 0.991 recall@10 at 0.23 ms in the selected run.

FAISSBEIRRay Data
280 Tesseract evaluations · VLMs pending Document AI evaluation

Build the OCR arena before publishing the leaderboard

A provider-neutral benchmark with exact synthetic truth, 13 seeded corruption families, durable page records, and one honestly scoped local baseline. The report separates real inference from adapter contract tests.

OCRVLMsModalRobustness
30 hand-labeled examples Evaluation

Calibrating the judge before trusting it

A deliberately small LLM-as-judge study with human labels, agreement metrics, and two CI gates: judge drift and answer-quality regression.

Cohen's κCI gatesClaude API
Synthetic data · planted bias Responsible decisions

Credit risk with a provable audit

An 800,000-application simulated portfolio with a known redlining proxy. The audit catches it and rejects the higher-Gini challenger.

LightGBMSHAPFairness
Dated, not “latest”

Frontier model decisions

Model names, access, licenses, and framework support change quickly. The playbook pins what was verified and defines how to rerun the decision.

Decision playbook · verified Jul 25, 2026

Kimi K3, Qwen3.6/3.7, and GLM-5.2

“Open” is not one category. Kimi K3 is currently an API model with weights announced for July 27; Qwen3.7 Max/Plus are hosted while Qwen3.6 has downloadable Apache-2.0 checkpoints; GLM-5.2 publishes MIT-licensed weights. The choice also depends on active parameters, context, multimodality, hardware, fine-tuning path, serving support, and quality per dollar.

Explicit design work

Decision products to build next

The architecture and evaluation are specified now; the status changes only after the app, data pipeline, and measurements exist.

Build dossier

Market intelligence

A sourced event timeline over filings, macro releases, news, prices, volume, and regime shifts with point-in-time storage and leakage-safe evaluation.

Read the design →
Build dossier

Personal health intelligence

Meals, labs, supplements, activity, sleep, family history, and genetics with consent, provenance, uncertainty, and culturally relevant evidence.

Read the safety-first design →
Production plan

Controllable AI film pipeline

A story and shot graph that joins ComfyUI generation with conventional edit, composite, color, sound, and continuity review.

Read the production plan →
Research map

Five years of AI

A source-linked timeline connecting models, robotics, world models, agents, video, audio, inference, and accelerators, and the ideas that endured.

Explore the timeline →