Terrane runtime v2.0

Run inference where your users are.

Terrane is a unified runtime for inference workloads, APIs and backend services. Region-aware routing, dynamic batching, elastic capacity and live observability, from a single deployment file.

runtime-status.log
Runtime live 24 regions / global edge

Control plane

Global deployment topology

Synced
P99 latency
18ms
Global traffic
14.2kreq/s
Regions online
24/24
Failover
none
Runtime health Stable

99.98%

Regional availability
Router events Live

    Failover scenario replays every 14 s

    Platform capabilities

    Infrastructure that ships with the model

    Deploy high-performance inference, orchestrate workloads automatically, route traffic globally and run on the frameworks your team already uses. The control plane underneath is ours to manage.

    Accelerated inference
    High throughput 8× H200 · 141GB VRAM Large-model inference · sustained traffic · low latency
    4× Gaudi 3 · 128GB VRAM Alternative accelerator tier · balanced loads
    2× L40S · 48GB VRAM

    Accelerated inference

    Run LLMs, embeddings and multimodal workloads on compute tiers sized for sustained production traffic, not demos.

    Elastic orchestration
    105 GPUs 50 GPUs 0 GPU
    Autoscaling Dynamically allocates active GPU capacity from live inference demand

    Serverless orchestration

    Scale from zero to multi-region capacity automatically, without tuning clusters, queues or hardware pools by hand.

    Global edge
    10 datacenters Paris16 ms Tokyo12 ms San Francisco21 ms Singapore21 ms
    $ terrane deploy llama-3 --regions sfo,tyo,par,sin

    Global edge inference

    Route every request to the healthiest nearby runtime so latency stays low and availability stays high, everywhere.

    Any framework
    PyTorchvLLMTensorRTONNXJAXCustom

    Any framework, any model

    PyTorch, TensorFlow, TensorRT, vLLM or your own container. Bring the stack you have, keep the deployment workflow.

    Core infrastructure

    Four modules, one runtime.

    Terrane brings deployment, routing, batching and model orchestration into one unified runtime so technical teams can ship production AI systems without stitching together fragmented infra.

    Module 01Active

    Global edge runtime

    Deploy inference workloads close to demand with region-aware routing, resilient failover and low-latency execution across the distributed runtime.

    Module 02ctrl.sys

    Model orchestration

    Manage model versions, hardware allocation, image builds and rollout policies from a single deployment layer designed for production inference.

    Module 03queue.act

    Dynamic batching

    Increase throughput automatically with request batching, queue-aware scheduling and runtime-level optimisations that improve GPU utilisation under live traffic.

    Module 04

    Unified deployment surface

    Ship APIs, inference endpoints and backend services through one deployment workflow with built-in routing, autoscaling, observability and infrastructure-aware execution.

    • Global routing
    • Autoscaling
    • Full observability
    terrane.deploy.yaml
    runtime: edge-inference
    regions: [iad, sfo, fra, sin]
    autoscale: enabled
    batching: dynamic
    observability: full
    Active regions4 / 4 online
    Median latency21ms
    Runtime statusHealthy / synchronised
    Last deploy synced 12 seconds ago
    Workflow

    How Terrane works

    From runtime setup to live observability, Terrane gives technical teams a clear operational path to deploy, route, scale and manage production AI workloads without stitching together fragmented infrastructure manually.

    1. Connect inputs#1

      Connect inputs.

      Bring models, APIs, queues and backend services into one deployment-ready control surface.

    2. Define runtime#2

      Define runtime.

      Configure compute, regions, scaling rules and rollout policies before traffic ever goes live.

    3. Deploy globally#3

      Deploy globally.

      Roll workloads out across regions through one workflow instead of managing isolated surfaces.

    4. Route traffic#4

      Route traffic.

      Send requests to the healthiest and closest runtime path using latency- and region-aware routing.

    5. Scale automatically#5

      Scale automatically.

      Expand active capacity under live inference demand without manually tuning queues or pools.

    6. Observe everything#6

      Observe everything.

      Track runtime health, request flow, regional status and system activity through one operational view.

    • 01

      Unified runtime

      Define deployment behaviour once and let Terrane manage execution across regions and hardware tiers.

    • 02

      Live observability

      Stay informed with real-time metrics for latency, health, throughput and regional runtime status.

    • 03

      Configurable policies

      Control compute classes, scaling thresholds, routing rules and rollout behaviour without rebuilding workflows.

    Proof / readiness

    Trusted for production workloads

    Terrane is built for teams running live inference, distributed APIs and regional runtime infrastructure, with reliability, observability and operational control designed into the platform.

    Operator feedback Platform team / applied AI
    “Terrane removed the infrastructure sprawl from our inference stack. We can deploy globally, monitor runtime health and scale traffic without building a custom control plane around it.”
    • Global routing
    • Autoscaling
    • Full observability
    Lead Platform Engineer Production inference team

    Performance visibility

    Monitor latency, throughput, queue health and regional runtime behaviour through one operational surface.

    Built for SRE teams · platform operators

    Runtime flexibility

    Deploy across regions, compute classes and model-serving stacks without rebuilding your delivery workflow.

    Supports vLLM · APIs · backend services

    Operational reliability

    Regional failover, health-aware routing and runtime safeguards keep production traffic stable under live demand.

    Used for Inference APIs · global runtimes

    Performance visibility

    Monitor latency, throughput, queue health and regional runtime behaviour through one operational surface.

    Built for SRE teams · platform operators

    Runtime flexibility

    Deploy across regions, compute classes and model-serving stacks without rebuilding your delivery workflow.

    Supports vLLM · APIs · backend services

    Operational reliability

    Regional failover, health-aware routing and runtime safeguards keep production traffic stable under live demand.

    Used for Inference APIs · global runtimes
    Availability
    0%

    Runtime health across active regions under sustained production traffic.

    Median latency
    0ms

    Region-aware request routing keeps inference performance close to end users.

    Active regions
    0

    Distributed runtime presence for global AI services and backend workloads.

    Scale profile
    0Burst

    Elastic runtime expansion aligned to live load, queue pressure and inference demand.

    Start deploying

    Ship to production this afternoon.

    Move from prototype to production with a unified runtime for inference, APIs and backend services, built for global routing, elastic scale and real operational visibility.

    • Global runtime
    • Elastic scaling
    • Built-in observability
    terrane.deploy
    Runtime profileedge-inference / global
    Live
    Regions24Active deployment regions
    Median latency21msGlobal request routing
    Deployment statusHealthy / synchronised
    • iad.runtimeOnline
    • fra.runtimeOnline
    • sin.runtimeOnline
    No cluster management required Built for production AI