Case study: a semantic search API at 14,000 requests per second
How a marketplace team moved embeddings, reranking and their public search API onto Terrane, cut p99 from 340 ms to 98 ms and stopped running a control plane.
How teams run inference APIs, backend services and batch workloads on Terrane in production.
Real workloads, real numbers, written with the teams who run them.
How a marketplace team moved embeddings, reranking and their public search API onto Terrane, cut p99 from 340 ms to 98 ms and stopped running a control plane.