Pack more models onto every chip — served through one API gateway.

OpenInfer oversubscribes your GPUs and CPUs — any mix of vendors and generations — to keep more models resident per device than naively fit, and serves them all through a unified, OpenAI-compatible gateway. Weave, our managed control plane, handles placement and routing across the fleet. We don’t scale your infrastructure; we plug into it — idle silicon becomes billable capacity.

One integrated system — control plane, routing, and serving.

Weave is a managed control plane. Add OpenInfer to the nodes you already run; each node dials out to Weave over one connection — no public IPs, no inbound ports — and reports what it can serve and how busy it is. Requests come down that connection to the best-fit node, and tokens stream back up. Routing and serving are built as one system, not an API gateway, a router, and a serving framework taped together.

Every request is a routing decision.

A load balancer sprays requests evenly. Weave reacts to live signals on every request — latency against each model’s SLA, where models are loaded, and what’s actually ready to serve.

01

SLA-aware rerouting

When a node’s latency drifts past its model’s SLA, Weave sheds new requests to a healthier node — before your customers feel it.

02

Model balancing & repacking

When demand for a model climbs, the control plane loads it onto more nodes — repacking the fleet to match the traffic, not a fixed layout.

03

Model-aware routing

Every request only lands on a node that has the model loaded, warmed, and with room to serve — never one that’s still downloading.

Pack more inference onto the hardware you already have.

OpenInfer treats a pile of mixed silicon as one dense pool of capacity — no standardizing your fleet, no rip-and-replace. More models per device, on whatever you run.

01

Oversubscribe every device

Pack many models onto each GPU or CPU. OpenInfer oversubscribes memory to keep more models resident per device than would naively fit — deploy and version them across the fleet from the Weave dashboard.

02

Any silicon, tuned per node

Mix GPU vendors, generations, and CPUs. The per-node runtime detects each device and tunes KV cache, weights, and memory for peak throughput — no per-node config.

Runs alongside Kubernetes, Ansible, CDK, or bare metal — add OpenInfer and nodes register and deregister automatically as your infra scales them. We never provision or scale machines; we plug into the infrastructure you own.

Sell to your customers, on shared silicon.

Carve one fleet across many tenants — with the isolation, controls, and accounting you need to run inference as a product. Every request is authenticated, scoped to an org, and metered before Weave ever picks a node.

Per-tenant keys

Issue scoped API keys per customer or project, and keep each tenant’s traffic isolated across the shared fleet.

Quotas & limits

Set quotas and rate limits per tenant, and decide whose requests get capacity first when the fleet is busy.

Usage metering

Per-tenant token and capacity accounting you can export and bill against.

Production-grade — and out of your data.

A managed control plane that keeps the fleet serving through failure and upgrades, and only ever sees what it needs to route.

Self-healing

Nodes stream their own health to Weave and count as present only while connected — so a failed node drops out instantly and cleanly, with no stale routing or manual re-registration.

Metadata-only routing

Weave routes on model and SLA metadata. It never inspects the content of your prompts or completions.

Don’t own the hardware? Use ours.

OpenInfer Cloud is the same platform, run by us on our Weave-managed fleet — a hosted, OpenAI-compatible API you can call today, no silicon required. It’s an on-ramp for teams without hardware, and proof the platform runs in production.

Deploy with the tooling you already use.

There’s no separate orchestration to run. The OpenInfer runtime deploys like any other workload — through the same infrastructure scaling toolkit you already use to bring GPU nodes up and down.

  1. 01

    Deploy the runtime

    Bake the OpenInfer runtime into your node image and roll it out with your existing scaling toolkit — Kubernetes, Ansible, CDK, or an autoscaler.

  2. 02

    The node dials Weave

    When a GPU node boots, the runtime opens a connection to the Weave router — no public IPs, no inbound ports.

  3. 03

    It packs models and comes online

    Weave sends the node its model-packing instructions; the runtime loads them and brings itself online, ready to serve.

Give it a try.

Bring OpenInfer to your hardware. Email us for access and we'll get you a node serving.