Ollama and the bundled model
Container deployments can include an Ollama model sidecar. With that local connector configured, inference stays on your infrastructure. Native/source installations need a separate Ollama service or another AI connector. Managed plans with vendor-hosted AI use DigitalOcean Gradient AI instead of a local sidecar. From 0.22.0, managed deployments do not offer a new Ollama configuration; existing records remain visible but cannot be edited there.
| Category | AI |
| Authentication | None |
| Reaches | Your Ollama host, by default the bundled one |
| Needs an agent | For private hosts unless the operator explicitly allows the destination; see below |
| Demo mode | Yes, off by default |
Capabilities
Section titled “Capabilities”ai.chat
Which powers the assistant, ticket triage, and the AI read on a ticket. See AI features.
What ships
Section titled “What ships”A local model service running Ollama, on the deployment’s own private network. The supplied Compose configuration does not publish its port through the public ingress. Keep that restriction when changing your stack.
The default is Qwen 2.5 3B Instruct at Q4 (qwen2.5:3b, about 2GB), chosen
to run on the class of machine schools buy, which is a small server
with no graphics card.
When LOCAL_AI_URL is configured, bootstrap creates the local connector for an
institution that has no AI connector. Existing or disabled connectors are preserved.
Managed vendor AI configuration takes precedence when supplied. Without a model
connector, the assistant uses its built-in command engine.
How fast it is
Section titled “How fast it is”Measured on a 4-core machine with no GPU, sharing itself with everything else on the box:
prompt 6.9s / 199 tokens generation 32s / 210 tokensSo 25 to 45 seconds for one answer, and generation is nearly all of it. Shorter prompts barely help. The two things that do are a shorter answer and better hardware.
A machine with a GPU does the same work five to ten times faster. If you have one, serve a model on it and point the OpenAI-compatible connector at it instead.
This is why the ticket summary loads itself when you open the page rather than sitting behind a button. On this hardware the wait has to happen somewhere, and the one place it must not happen is in front of somebody who has just clicked.
Requirements
Section titled “Requirements”| Disk | About 2GB for the default model |
| Memory | 4GB free, on top of everything else |
| GPU | Not required, and it makes a large difference if present |
The model is fetched on first boot, so an update is not also a two gigabyte download of unchanged weights.
Configuring it
Section titled “Configuring it”Admin, Connectors, Ollama (local AI), Configure.
| Field | Default | Value |
|---|---|---|
baseUrl |
http://localhost:11434 |
The Ollama host. The bundled one is http://local-ai:11434 |
model |
qwen2.5:3b |
Any model pulled on that host |
timeoutMs |
60000 |
Give up on an answer after this long |
maxTokens |
768 |
The most the model may write in one answer |
demoMode |
false |
Skip the model and use the built-in command engine only |
There are no credentials.
Private network access
Section titled “Private network access”Ollama follows the outbound destination policy. A private host needs a
connector agent, the
operator’s ALLOW_PRIVATE_EGRESS=1 setting, or an origin matching the operator’s
LOCAL_AI_URL exactly (scheme, hostname and port). The last option permits the
configured local sidecar without opening other private hosts to connector settings.
Cloud metadata destinations remain blocked, including hostname aliases.
Changing a connector’s baseUrl does not change the operator’s allowed origin.
Public destinations remain subject to the normal outbound checks.
The supplied Compose stack pins Ollama to 0.32.14. Plugboard adapts large schema bounds that this version refuses and validates the returned structured answer. If you override the image, test a ticket brief after the change; a healthy connection test alone does not prove structured answers work.
Swapping the model
Section titled “Swapping the model”If you have better hardware, pull a larger model on the Ollama host and change
model. Nothing else changes.
LOCAL_AI_MODEL decides what gets pulled, so you change the model rather than
pulling a second one alongside it.
One call at a time
Section titled “One call at a time”The connector serialises calls to a host and gives up on one after the timeout without retrying it.
A machine with no GPU serves one completion at a time. Overlapping requests used to turn into several timeouts, where queuing them produces several answers, slowly.
Managed hosting does not run this
Section titled “Managed hosting does not run this”Region hosts do not start the bundled model. A shared machine serving several schools cannot also hold a resident model without taking memory from their databases.
On managed hosting, AI comes either from your own key through Anthropic or an OpenAI-compatible endpoint, or from us if your plan includes it. See three ways to get a model.
Where inference runs
Section titled “Where inference runs”Inference goes to the configured baseUrl. A local sidecar keeps those model
requests on the deployment’s network; pointing the connector at an external
Ollama host sends them there. Other integrations and MCP clients have their own
data flows. The compliance page shows the configured
AI provider.
Troubleshooting
Section titled “Troubleshooting”| Symptom | Cause |
|---|---|
| The assistant only handles exact phrases | demoMode is on, or no AI connector is enabled |
| Connection refused | The bundled service is not running, or baseUrl is wrong. Inside the stack it is http://local-ai:11434 |
| Everything times out | The machine is too slow for the model. Raise timeoutMs, use a smaller model, or move to a GPU |
| Blocked by the egress guard | Check the agent binding, operator private-egress setting or exact LOCAL_AI_URL origin. Metadata addresses remain blocked |
model not found |
Not pulled on that host. Run ollama pull |
| Answers are confidently wrong | A 3B model is small. The confirmation prompt before any change is why this is annoying rather than dangerous |