The runtime layer is provider-neutral. Cloud services remain primary, optional Ollama models can participate where deployed, and user-controlled GPU endpoints can register as compatible runtimes without making Agent Forge itself GPU-dependent.
The fleet is discovered from configured providers and compatible runtimes. The website does not promise a fixed model count or permanent roster. Three logical core slots organize primary work; a broader pool can support agents, subtasks, and fallback candidates.
Agent Forge can use hosted APIs, Ollama-backed cloud or local models, and user-owned compatible endpoints. Open-weight components may support embeddings or classification, but the platform is not limited to Hugging Face models.
A user-owned GPU runtime can participate through an authenticated compatible endpoint. Registration, credentials, model metadata, reachability, and ongoing heartbeats are required. GPU connectivity is independent of the normal cloud-first path.