Xinity Runtime is a self-hosted, open-source AI platform for on-premise LLM inference. It exposes an OpenAI-compatible API with multi-model orchestration and a management dashboard, letting regulated organizations run generative AI on their own hardware with zero data egress.