---
title: Local models
group: Apps
order: 5
summary: Use the desktop Ollama library, manage downloads and memory, and test a model in the playground.
updated: 2026-10-06
---
Local models run through Ollama on your computer. The desktop library can be browsed without a running engine; downloading, loading and generation need Ollama. Local models are separate from the website's hosted model selection and from any separately configured shared model pool.

## Prerequisites

Install the current supported [desktop package](/#download), then install/start Ollama using its own instructions. You need internet for model downloads, disk space for the chosen weights and enough available memory to run them. The library's fit estimate is a guide; it does not guarantee performance or GPU compatibility.

The implementation talks to Ollama's local API. Do not expose that API publicly to let the desktop app reach it. A remote/shared inference endpoint needs its own authorization and deployment setup.

## First model

1. Open **AI > Local models** and confirm engine status.
2. Choose a small model/variant whose download and memory requirements fit.
3. Start the download and inspect its progress. You can browse while it runs.
4. Open the installed list and load it if needed.
5. In the playground, select the model and ask a short factual/local test.

The installed model appears and its reply streams in the playground. A finished download does not show that the model can load into available memory or use tools correctly.

## Manage models

The library includes model metadata, installed/running lists, pull, show, chat/generate, unload and remove operations. Loading holds weights in memory; unloading releases them. Removing deletes the local model and is a destructive action subject to registry approval. Cancelling a download stops that request; inspect the engine's remaining files/state before retrying.

The model catalog is curated, and capability labels depend on the model. A text-only model cannot accept images, even if another variant in the family can. Review license and variant details before using it for a project.

## Assistant and voice

Desktop voice can be configured to use a local model for planning. Hosted assistant routing and local model pools are separate features. Choosing a local playground model does not automatically make every integration task or speech-recognition step offline. Calls to cloud services still need a connection and their own credentials.

## Recovery

| Symptom | Recovery |
|---|---|
| “Ollama is not running” | Start the engine, check its local endpoint, then retry. |
| Download stalls/fails | Check disk/network and the engine error, then retry. |
| Model cannot load | Unload other models/apps, select a smaller variant and check memory. |
| Slow response | Try a smaller model/context and inspect engine CPU/GPU use. |
| Model invents a tool/result | Verify real activity; select a suitable tool-capable model or use hosted routing. |

A fluent answer does not mean a command ran. Inspect capability/audit results for actual operations. Related: [Desktop](/docs/desktop), [Assistant](/docs/assistant), [Voice](/docs/voice).
