Starlight
AI inference
Private AI inference for the enterprise.
Serve LLMs on infrastructure you control, with the same security, governance, and audit posture as the rest of your estate.
Your models. Your hardware. Your data.
Enterprise LLM serving with the security posture of the rest of your estate, on hardware you own.
OpenAI-compatible endpoints. On infrastructure you control.
Models packaged and pulled as OCI artifacts, from your registries.
vLLM and llama.cpp backends behind one endpoint shape.
Runs connected, degraded, or fully air-gapped.
One identity, policy, and audit surface across VMs, containers, and AI.
AI supply chain
Models delivered as OCI containers.
AI models are infrastructure artifacts, not loose blobs. Pull them from a public or private registry the same way you pull any container image. The runtime starts, the endpoint comes up.
Give AI models the same enterprise treatment as container workloads.
Inference backends
vLLM for production GPU. llama.cpp for constrained hosts.
Two proven open inference backends behind one endpoint shape. Applications written against cloud inference APIs run unchanged against a local Starlight endpoint.
One endpoint shape. The platform picks the backend; the client code does not change.
Hardware discovery
One model. Heterogeneous hardware.
The platform handles hardware discovery, so a single model artifact serves every class of host in your fleet, from GPU servers to CPU-only sites.
Ship AI services faster, on hardware you already own.
OS foundation
Drivers in the image. Atomic updates.
AI deployments fail on driver pairing: kernel and driver updated independently, container toolkit out of sync, CDI broken on the next reboot. Starlight's hardened OS removes that failure class.
GPU drivers and kernel stay in sync and signed.
Hardware security
Confidential compute. Hardware-isolated inference.
Inference over sensitive data deserves protection while it runs. Memory is protected while the workload is running, not just at rest or in transit.
Workload runtime
Built-in runtime protection across VMs, containers, and AI.
The same enforcement layer applies across the entire workload surface, inference included. It ships with the platform, with no separate agent install.
Inference your security team can sign off on.
Network policy
Declarative network intent for AI endpoints.
Describe intent once. The platform compiles L3 and L4 firewall rules and L7 gateway configuration from the same artifact, so inference endpoints are reachable by exactly what you intend and nothing else.
Governance
MCP and agents. Same identity. Same policy. Same audit.
Whether a human, a pipeline, or an AI agent is driving the platform, it is the same API surface under one identity model. For sandboxing the agents themselves, see Agent microVMs.
AI security
Starlight AI-SPM.
Posture management across models, datasets, agents, and pipelines, mapped to the frameworks your auditors already use.
Detect and block AI-specific threats via model red-teaming, prompt filtering, and secure ML supply-chain controls.
AI you can govern, audit, and defend.
Resilience
Inference that keeps working when the network doesn't.
Local inference means availability does not depend on a connection. Disconnected, intermittent, and limited bandwidth are first-class operating conditions, and local workloads keep running across all three.
The host is the source of truth. Connectivity is a feature, not a dependency.
Enterprise AI, inside your boundary.
Bring inference to your infrastructure.
Tell us about your environment and your models, and we will walk through what private inference on Starlight looks like.
