ch 01 · the spark
every build starts with
a what if?
Turning them into working software is the job. Keeping them honest is the craft. This channel is a tour of the receipts — five minutes, no vibes.
ch 02 · the gap
fluent ≠ correct.
100%
grounded baseline · pass rate + root-cause accuracy
10%
Qwen 2.5 3B · 90% average confidence
TraceBack, 30 runs per configuration × 3 scenarios. The most dangerous failures sound confident. So everything I build gets measured against deterministic ground truth.
ch 03 · the pipeline
promotion is earned.
register → version → gate → serve → watch
ModelDock won't promote a model that didn't pass the eval-threshold gate. 72 tests in CI/CD, authenticated inference, drift monitoring. A registry that says no.
ch 04 · the receipts
don't take my word for it.
- dasaiko +86% recall@5 · live at dasaiko.dev ↗
- traceback 100% grounded vs 10% @ 90% conf ↗
- modeldock ★9 · 72 tests · gated lifecycle ↗
- provena #153 ✓ 498 passed · 33 skipped ↗
- clearlabel-ai yolov8 auto-annotation · dual audit ↗
full list on the brief →
ch 05 · the signal
still curious, out loud.
more on linkedin ↗
end of broadcast