http://localhost:11434 via OLLAMA_BASE_URL in backend configuration).
Curated models
When the Ollama host is reachable, live discovery can surface additional chat models beyond this curated list. Runtime may also accept extra embeddings (for example
nomic-embed-text, all-minilm, or Jina EN) that are not in the picker — see Providers.
Why teams use this
- No model traffic leaves your servers when inference is fully local
- Aligns with sovereignty positioning: self-hosted / air-gapped, DSGVO-minded deployments
- Works with the same AI Search / AI Chatbot / widget surfaces as cloud providers
Practical setup
- Install and run Ollama on a host reachable from the RAGSuite API.
- Set
OLLAMA_BASE_URLin backend env. - Select Custom LLM / Ollama and a curated (or discovered) model in the console for the project.
- Verify answers still cite your ingested sources (retrieval remains yours either way).