Market intelligence
A sourced event timeline over filings, macro releases, news, prices, volume, and regime shifts with point-in-time storage and leakage-safe evaluation.
Read the design →I work across data curation, full and parameter-efficient fine-tuning, reinforcement learning, retrieval, multimodal inference, evaluation, and the decision interface around the model.
Results include the regressions, domain shifts, and small-sample limits that change what the headline means.
A license-gated 14,080-pair corpus, full BF16 FSDP2 training on four H100s, FP8 export and vLLM serving, then GRPO with the compiler as verifier. Domain exact match reached 23%, while a 7.3-point HumanEval regression tripped the retention gate.
Clean + real-meeting evaluation AudioWER and DER implemented from scratch, held-out calibration, a browser-tab streaming path, and the honest result that the clean-benchmark ranking reverses on AMI meetings.
Six embedders · five indexes RetrievalEmbedding models are judged against relevance; FAISS indexes are judged against exact neighbors. HNSW reaches 0.991 recall@10 at 0.23 ms in the selected run.
280 Tesseract evaluations · VLMs pending Document AI evaluationA provider-neutral benchmark with exact synthetic truth, 13 seeded corruption families, durable page records, and one honestly scoped local baseline. The report separates real inference from adapter contract tests.
30 hand-labeled examples EvaluationA deliberately small LLM-as-judge study with human labels, agreement metrics, and two CI gates: judge drift and answer-quality regression.
Synthetic data · planted bias Responsible decisionsAn 800,000-application simulated portfolio with a known redlining proxy. The audit catches it and rejects the higher-Gini challenger.
Model names, access, licenses, and framework support change quickly. The playbook pins what was verified and defines how to rerun the decision.
“Open” is not one category. Kimi K3 is currently an API model with weights announced for July 27; Qwen3.7 Max/Plus are hosted while Qwen3.6 has downloadable Apache-2.0 checkpoints; GLM-5.2 publishes MIT-licensed weights. The choice also depends on active parameters, context, multimodality, hardware, fine-tuning path, serving support, and quality per dollar.
The architecture and evaluation are specified now; the status changes only after the app, data pipeline, and measurements exist.
A sourced event timeline over filings, macro releases, news, prices, volume, and regime shifts with point-in-time storage and leakage-safe evaluation.
Read the design →Meals, labs, supplements, activity, sleep, family history, and genetics with consent, provenance, uncertainty, and culturally relevant evidence.
Read the safety-first design →A story and shot graph that joins ComfyUI generation with conventional edit, composite, color, sound, and continuity review.
Read the production plan →A source-linked timeline connecting models, robotics, world models, agents, video, audio, inference, and accelerators, and the ideas that endured.
Explore the timeline →A smaller doorway into the broader reference library.
VLA models, diffusion policies, world models, simulation, transfer, and evaluation.
Prefill/decode separation, batching, KV caches, parallelism, observability, and cost.
Bayer demosaicing, convolution, edge-aware smoothing, camera controls, and runnable C99 and OpenCV C++.
Concrete reduced-scale builds across language, vision, audio, RL, robotics, retrieval, and systems.
From value learning and policy gradients through PPO, GRPO, self-play, and verifiable rewards.