SLA-aware rerouting
When a node’s latency drifts past its model’s SLA, Weave sheds new requests to a healthier node — before your customers feel it.
OpenInfer oversubscribes your GPUs and CPUs — any mix of vendors and generations — to keep more models resident per device than naively fit, and serves them all through a unified, OpenAI-compatible gateway. Weave, our managed control plane, handles placement and routing across the fleet. We don’t scale your infrastructure; we plug into it — idle silicon becomes billable capacity.
Weave is a managed control plane. Add OpenInfer to the nodes you already run; each node dials out to Weave over one connection — no public IPs, no inbound ports — and reports what it can serve and how busy it is. Requests come down that connection to the best-fit node, and tokens stream back up. Routing and serving are built as one system, not an API gateway, a router, and a serving framework taped together.
A load balancer sprays requests evenly. Weave reacts to live signals on every request — latency against each model’s SLA, where models are loaded, and what’s actually ready to serve.
When a node’s latency drifts past its model’s SLA, Weave sheds new requests to a healthier node — before your customers feel it.
When demand for a model climbs, the control plane loads it onto more nodes — repacking the fleet to match the traffic, not a fixed layout.
Every request only lands on a node that has the model loaded, warmed, and with room to serve — never one that’s still downloading.
OpenInfer treats a pile of mixed silicon as one dense pool of capacity — no standardizing your fleet, no rip-and-replace. More models per device, on whatever you run.
Pack many models onto each GPU or CPU. OpenInfer oversubscribes memory to keep more models resident per device than would naively fit — deploy and version them across the fleet from the Weave dashboard.
Mix GPU vendors, generations, and CPUs. The per-node runtime detects each device and tunes KV cache, weights, and memory for peak throughput — no per-node config.
Runs alongside Kubernetes, Ansible, CDK, or bare metal — add OpenInfer and nodes register and deregister automatically as your infra scales them. We never provision or scale machines; we plug into the infrastructure you own.
Carve one fleet across many tenants — with the isolation, controls, and accounting you need to run inference as a product. Every request is authenticated, scoped to an org, and metered before Weave ever picks a node.
Issue scoped API keys per customer or project, and keep each tenant’s traffic isolated across the shared fleet.
Set quotas and rate limits per tenant, and decide whose requests get capacity first when the fleet is busy.
Per-tenant token and capacity accounting you can export and bill against.
A managed control plane that keeps the fleet serving through failure and upgrades, and only ever sees what it needs to route.
Nodes stream their own health to Weave and count as present only while connected — so a failed node drops out instantly and cleanly, with no stale routing or manual re-registration.
Weave routes on model and SLA metadata. It never inspects the content of your prompts or completions.
OpenInfer Cloud is the same platform, run by us on our Weave-managed fleet — a hosted, OpenAI-compatible API you can call today, no silicon required. It’s an on-ramp for teams without hardware, and proof the platform runs in production.
There’s no separate orchestration to run. The OpenInfer runtime deploys like any other workload — through the same infrastructure scaling toolkit you already use to bring GPU nodes up and down.
Bake the OpenInfer runtime into your node image and roll it out with your existing scaling toolkit — Kubernetes, Ansible, CDK, or an autoscaler.
When a GPU node boots, the runtime opens a connection to the Weave router — no public IPs, no inbound ports.
Weave sends the node its model-packing instructions; the runtime loads them and brings itself online, ready to serve.
Bring OpenInfer to your hardware. Email us for access and we'll get you a node serving.