About

A runtime team, not a platform company.

Terrane exists because production AI is not a model behind a URL. It is a model, an API in front of it, workers around it and a control plane nobody wanted to write. We wrote it so you do not have to.

Mission

Keep production traffic flowing.

When hardware fails, when a region goes dark, when a deploy goes wrong. That is the whole job, and it is harder than it sounds. Most of what we have built in three years is invisible when it works: routing that sheds load before anyone notices, rollouts that pause themselves, batching that respects a latency budget under a traffic spike.

We sell that invisibility to teams who would rather spend their engineers on the model than on the plumbing around it. Our customers describe the outcome the same way, almost every time: "we had a search team and an infrastructure team; now we have a search team."

Founded
2023
Based in
San Francisco, California
Regions operated
24
Team
38 people, 29 engineers
Principles

How we decide what to build

  1. 01

    One thing, reliably

    We run production traffic. We do not host training, sell a vector database or ship an agent framework. Every feature has to make the runtime more dependable or it does not ship.

  2. 02

    Explainable in one sentence

    The routing score is a sum you can read. The batching control is one number. If an operator cannot predict what the system will do at 3 a.m., the design is wrong.

  3. 03

    Failure is a ramp, not a cliff

    Regions degrade gradually, rollouts pause instead of exploding, capacity drains before it disappears. Sudden state changes are where incidents come from.

  4. 04

    Numbers over adjectives

    We publish p99s, not promises. Every claim on this site has a measurement behind it and most of them have a blog post.

Timeline

Three years, two runtimes.

Every step on the way to 2.0 was a system we ran in production first and productised second.

  1. 2023

    Terrane is founded by three engineers who had each built an inference control plane by hand and did not want to do it a third time.

  2. 2024

    Runtime 1.0 ships with region-aware routing across eight regions. First production customers move embedding and reranking workloads.

  3. 2025

    Dynamic batching replaces static batch configuration. The fleet grows to eighteen regions with uniform hardware tiers.

  4. 2026

    Runtime 2.0 unifies inference, HTTP services and workers under one deployment surface. 24 regions, rebuilt observability.

Work with us

Run your stack on Terrane.