Dynamic batching: more throughput without wrecking the tail
Batching requests is the cheapest way to raise GPU utilisation and the easiest way to ruin latency. Here is how Terrane decides when to wait and when to fire.
Product Lead, Observability
Product lead for observability and batching. Believes every dashboard should answer a question in under three seconds.
Lin owns the parts of Terrane you look at when something feels off: runtime health, request flow, regional status and the batching controls that trade latency for throughput. She previously led the metrics product at an APM vendor and spent two years as an SRE on a payments platform.
Her writing focuses on how operators actually reason about live systems, and how the product can get out of their way.
2 articles in the journal.
Batching requests is the cheapest way to raise GPU utilisation and the easiest way to ruin latency. Here is how Terrane decides when to wait and when to fire.
How a marketplace team moved embeddings, reranking and their public search API onto Terrane, cut p99 from 340 ms to 98 ms and stopped running a control plane.