Skip to content

Changelog

What shipped recently

Product updates and SDK releases. Subscribe via RSS or watch the GitHub repo.

August 5, 2026 — Gradual prompt rollout

You can now send a percentage of traffic to a new prompt revision instead of cutting over all at once. Set the split in the prompt editor, watch the eval score on each arm, then ramp or revert. Available on Team and above.

July 28, 2026 — Judge calibration reports

Every LLM judge now reports how often it agrees with human labels on the same examples. Judges below 80% agreement are flagged in the UI, because gating a deploy on an uncalibrated judge is worse than not gating at all.

July 14, 2026 — Trace search rewrite

Full-text search across prompts and completions is roughly 12x faster and now supports boolean operators and field-scoped queries. Searches that used to time out on large workspaces now return in under a second.

June 30, 2026 — Anthropic extended thinking support

The Anthropic integration now captures extended thinking blocks as a separate span, so you can inspect the reasoning that led to a response without it cluttering the completion view.

June 18, 2026 — Dataset splits

Datasets support train, dev and holdout splits. Evaluations can target a specific split, which makes it possible to keep a genuine holdout set that your prompt iteration never sees.

June 2, 2026 — Self-hosted Terraform module

Self-hosted deployment now ships as a Terraform module alongside the existing Helm chart. Provisions the full stack including Postgres, object storage and workers.

May 19, 2026 — Cost anomaly alerts

New monitor type that alerts on token spend deviating from its trailing baseline, broken down by feature, model and prompt version.

Stop guessing whether your AI feature got better

Free for 50,000 traces a month. No credit card, no sales call, five minutes to your first trace.

Start free