A runtime team, not a platform company.
Terrane exists because production AI is not a model behind a URL. It is a model, an API in front of it, workers around it and a control plane nobody wanted to write. We wrote it so you do not have to.
Keep production traffic flowing.
When hardware fails, when a region goes dark, when a deploy goes wrong. That is the whole job, and it is harder than it sounds. Most of what we have built in three years is invisible when it works: routing that sheds load before anyone notices, rollouts that pause themselves, batching that respects a latency budget under a traffic spike.
We sell that invisibility to teams who would rather spend their engineers on the model than on the plumbing around it. Our customers describe the outcome the same way, almost every time: "we had a search team and an infrastructure team; now we have a search team."
- Founded
- 2023
- Based in
- San Francisco, California
- Regions operated
- 24
- Team
- 38 people, 29 engineers
How we decide what to build
- 01
One thing, reliably
We run production traffic. We do not host training, sell a vector database or ship an agent framework. Every feature has to make the runtime more dependable or it does not ship.
- 02
Explainable in one sentence
The routing score is a sum you can read. The batching control is one number. If an operator cannot predict what the system will do at 3 a.m., the design is wrong.
- 03
Failure is a ramp, not a cliff
Regions degrade gradually, rollouts pause instead of exploding, capacity drains before it disappears. Sudden state changes are where incidents come from.
- 04
Numbers over adjectives
We publish p99s, not promises. Every claim on this site has a measurement behind it and most of them have a blog post.
Three years, two runtimes.
Every step on the way to 2.0 was a system we ran in production first and productised second.
- 2023
Terrane is founded by three engineers who had each built an inference control plane by hand and did not want to do it a third time.
- 2024
Runtime 1.0 ships with region-aware routing across eight regions. First production customers move embedding and reranking workloads.
- 2025
Dynamic batching replaces static batch configuration. The fleet grows to eighteen regions with uniform hardware tiers.
- 2026
Runtime 2.0 unifies inference, HTTP services and workers under one deployment surface. 24 regions, rebuilt observability.
The people behind the journal
Engineers and product leads who run the runtime and write about it. Every article carries a byline you can follow.
- DR Daniel Reyes Developer Advocate Developer Advocate. Writes the tutorials he wishes existed when he was on call. Articles
- LZ Lin Zhao Product Lead, Observability Product lead for observability and batching. Believes every dashboard should answer a question in under three seconds. Articles
- MO Maya Okonkwo Head of Runtime Engineering Head of Runtime Engineering. Spent a decade making distributed systems boring on purpose. Articles