Shipping gradual prompt rollout
How we built percentage-based traffic splitting for prompt revisions, and the design decisions that were harder than expected.
Sofia Lindqvist ·
Engineering notes, evaluation techniques and product updates from the team building Cloudmind.
How we built percentage-based traffic splitting for prompt revisions, and the design decisions that were harder than expected.
Sofia Lindqvist ·
A practical method for going from zero to a 50-example golden dataset that actually catches regressions, using traces you already have.
Priya Nandakumar ·
LLM judges are useful and widely misused. Here is how to measure whether yours is good enough to gate a deploy on.
Marcus Hale ·
A walkthrough of how a silent retry loop tripled one team's token bill, and how nested tracing made it obvious in ten minutes.
Sofia Lindqvist ·
The case for moving prompts out of source control and into a versioned registry, and the objections that are actually right.
Marcus Hale ·
Retrieval quality, not generation quality, is where most RAG systems fail. Here is how to measure the two separately.
Priya Nandakumar ·
Traditional APM assumes failures are loud. LLM failures are quiet, well-formed and return 200 OK. That difference breaks most of the tooling.
Marcus Hale ·
Announcing our Series A, led by Ridgeline Ventures, and what we are building next.
Cloudmind Team ·