InferHub
Self-hosted LLM inference mesh in .NET. One Ollama-compatible API in front, a pool of GPU worker nodes behind it — run the hub where you have no GPU, run nodes where you do. Pluggable backends (Ollama first).
Self-hosted LLM inference mesh in .NET. One Ollama-compatible API in front, a pool of GPU worker nodes behind it — run the hub where you have no GPU, run nodes where you do. Pluggable backends (Ollama first).