Skip to content

Ollama and the bundled model

Container deployments can include an Ollama model sidecar. With that local connector configured, inference stays on your infrastructure. Native/source installations need a separate Ollama service or another AI connector. Managed plans with vendor-hosted AI use DigitalOcean Gradient AI instead of a local sidecar. From 0.22.0, managed deployments do not offer a new Ollama configuration; existing records remain visible but cannot be edited there.

Category AI
Authentication None
Reaches Your Ollama host, by default the bundled one
Needs an agent For private hosts unless the operator explicitly allows the destination; see below
Demo mode Yes, off by default

ai.chat

Which powers the assistant, ticket triage, and the AI read on a ticket. See AI features.

A local model service running Ollama, on the deployment’s own private network. The supplied Compose configuration does not publish its port through the public ingress. Keep that restriction when changing your stack.

The default is Qwen 2.5 3B Instruct at Q4 (qwen2.5:3b, about 2GB), chosen to run on the class of machine schools buy, which is a small server with no graphics card.

When LOCAL_AI_URL is configured, bootstrap creates the local connector for an institution that has no AI connector. Existing or disabled connectors are preserved. Managed vendor AI configuration takes precedence when supplied. Without a model connector, the assistant uses its built-in command engine.

Measured on a 4-core machine with no GPU, sharing itself with everything else on the box:

prompt 6.9s / 199 tokens generation 32s / 210 tokens

So 25 to 45 seconds for one answer, and generation is nearly all of it. Shorter prompts barely help. The two things that do are a shorter answer and better hardware.

A machine with a GPU does the same work five to ten times faster. If you have one, serve a model on it and point the OpenAI-compatible connector at it instead.

This is why the ticket summary loads itself when you open the page rather than sitting behind a button. On this hardware the wait has to happen somewhere, and the one place it must not happen is in front of somebody who has just clicked.

Disk About 2GB for the default model
Memory 4GB free, on top of everything else
GPU Not required, and it makes a large difference if present

The model is fetched on first boot, so an update is not also a two gigabyte download of unchanged weights.

Admin, Connectors, Ollama (local AI), Configure.

Field Default Value
baseUrl http://localhost:11434 The Ollama host. The bundled one is http://local-ai:11434
model qwen2.5:3b Any model pulled on that host
timeoutMs 60000 Give up on an answer after this long
maxTokens 768 The most the model may write in one answer
demoMode false Skip the model and use the built-in command engine only

There are no credentials.

Ollama follows the outbound destination policy. A private host needs a connector agent, the operator’s ALLOW_PRIVATE_EGRESS=1 setting, or an origin matching the operator’s LOCAL_AI_URL exactly (scheme, hostname and port). The last option permits the configured local sidecar without opening other private hosts to connector settings. Cloud metadata destinations remain blocked, including hostname aliases.

Changing a connector’s baseUrl does not change the operator’s allowed origin. Public destinations remain subject to the normal outbound checks.

The supplied Compose stack pins Ollama to 0.32.14. Plugboard adapts large schema bounds that this version refuses and validates the returned structured answer. If you override the image, test a ticket brief after the change; a healthy connection test alone does not prove structured answers work.

If you have better hardware, pull a larger model on the Ollama host and change model. Nothing else changes.

LOCAL_AI_MODEL decides what gets pulled, so you change the model rather than pulling a second one alongside it.

The connector serialises calls to a host and gives up on one after the timeout without retrying it.

A machine with no GPU serves one completion at a time. Overlapping requests used to turn into several timeouts, where queuing them produces several answers, slowly.

Region hosts do not start the bundled model. A shared machine serving several schools cannot also hold a resident model without taking memory from their databases.

On managed hosting, AI comes either from your own key through Anthropic or an OpenAI-compatible endpoint, or from us if your plan includes it. See three ways to get a model.

Inference goes to the configured baseUrl. A local sidecar keeps those model requests on the deployment’s network; pointing the connector at an external Ollama host sends them there. Other integrations and MCP clients have their own data flows. The compliance page shows the configured AI provider.

Symptom Cause
The assistant only handles exact phrases demoMode is on, or no AI connector is enabled
Connection refused The bundled service is not running, or baseUrl is wrong. Inside the stack it is http://local-ai:11434
Everything times out The machine is too slow for the model. Raise timeoutMs, use a smaller model, or move to a GPU
Blocked by the egress guard Check the agent binding, operator private-egress setting or exact LOCAL_AI_URL origin. Metadata addresses remain blocked
model not found Not pulled on that host. Run ollama pull
Answers are confidently wrong A 3B model is small. The confirmation prompt before any change is why this is annoying rather than dangerous