Ollama 0.40.2 Upgrades Local Models in the Background and Keeps the Old Copies as Backups
Ollama v0.40.2 upgrades models pulled on earlier versions to llama.cpp in the background the first time they run, and keeps the originals as backups. What that means for disk budgets, safe downgrades, canarying local fleets and the documented jq cleanup loop.
Ollama shipped v0.40.2 on October 8, a day after v0.40.1, and the small release changes something operators feel more than any feature line: custody of the models already on disk. Models downloaded with earlier versions are upgraded in the background the first time you run them, for better performance and compatibility when running on llama.cpp, and Ollama keeps the original copy as a backup so downgrading stays safe. The releases index marks v0.40.2 as the latest release; the same index dates v0.40.0 to Sep 25, the release in which supported model architectures began running on MLX by default on Apple Silicon devices.
If you run local models for coding agents, overnight jobs or a self-hosted workstation, this quiet upgrade is worth planning around: it moves your disk budget, it moves the downgrade story, and it changes what ‘the same model’ means on a host.
What happens on first run
The v0.40.2 notes describe a per-model migration triggered by use, not by the installer:
- You run a model pulled on an earlier Ollama, and it is upgraded in the background to run on llama.cpp.
- The original copy is kept as a backup, so upgraded models are temporarily kept on disk.
- A future release will remove these backups automatically; until then, deletion is opt-in.
ollama listno longer shows duplicate entries for upgraded models (fixed in pull request 18874), so inventory scripts see one entry per model again.

The disk bill, and the documented way out
While a backup exists, both the upgraded copy and the original occupy space, and the release notes publish no sizes for either, so the only honest number is the one you measure: record your model-directory size before upgrading, run each model once, then record it again. Hosts with small disks, quotas or network volumes are the ones that break.
If you need the space back now, the release notes ship a documented cleanup loop (requires jq) that selects ggml backup digests for models whose manifests include a llamacpp runner and feeds them to ollama rm:
curl -s localhost:11434/api/tags | jq -r '.models[].name' | while read -r m; do
curl -s localhost:11434/api/show -d "{\"model\":\"$m\"}" | jq -r \
'select(any(.manifests[]?; .runner == "llamacpp")) | .manifests[] | select(.runner == "ggml") | .digest'
done | sort -u | xargs -n1 ollama rm
Two caveats, both from the notes: the loop deletes only the backups, and if you later downgrade to a version older than 0.40 you will need to re-pull those models. That turns backup retention into a policy decision, not housekeeping:
The rest of v0.40.2 and v0.40.1
| Change | Release | Why it matters |
|---|---|---|
ollama launch claude uses the model’s full context length (18855) |
v0.40.2 | The launch command itself now runs at the model’s full context length; downstream effects belong in your canary checks |
ollama list duplicate-entry fix (18874) |
v0.40.2 | Fleet inventories stop double-counting upgraded models |
| Server proxies cloud usage and balance APIs (18829) | v0.40.1 | The local server proxies the cloud usage and balance APIs, one endpoint for hybrid setups |
| Clef head-read fix past 2GiB on Windows (18777) | v0.40.1 | Shipped commit line: clef head reads past 2GiB on Windows are fixed; verify on your Windows hosts |
| Manifests avoid symlinks on Windows (18852) | v0.40.1 | Shipped commit line: Windows manifests avoid symlinks; verify on your Windows volumes |
| CLI onboarding drops the account step (18826) | v0.40.1 | Onboarding no longer includes the account step, per the commit line |
| MLX drops the carried Metal residency patch (18854) | v0.40.1 | The patch is upstream now, so Ollama drops its carried copy |
The repository README shows why the ollama launch claude line matters beyond Ollama’s own CLI: supported coding integrations include Claude Code, Codex, Copilot CLI, DeepSeek Harness, Droid and OpenCode, so this local runtime sits under a fair amount of agent tooling.
Why this deserves a canary
A background model upgrade on first run is exactly the kind of default drift that Canary Every CLI Upgrade warns about: the same model name now resolves to a different binary on disk, with no operator action and no pinned version. Before rolling v0.40.2 across a fleet:
- Upgrade one host first, run the models your agents actually load, and diff behavior against your eval set.
- Check on the canary whether the full-context behavior of
ollama launch claudeshifts compaction or token spend; pricing self-hosted runs per completed task, as in ‘Free’ Weights, Paid Tasks, matters double when the context window under an agent changes. - Decide the backup policy per host class: a local AI workstation can purge backups after verification, while hosts with a downgrade plan should keep them until 0.40 proves stable.
What to do now
- Check disk headroom before upgrading hosts with large model libraries; assume both copies coexist until you act.
- Run each model once on a canary host so the upgrade happens on purpose, not mid-job overnight.
- Pick a backup policy: keep, purge now with the documented loop, or wait for automatic removal in a future release.
- Windows fleets should take v0.40.1 seriously: its commit lines fix clef head reads past 2GiB on Windows and make manifests avoid symlinks there; verify both on your Windows hosts.
- Re-run inventory after upgrading:
ollama listis fixed, and scripts that compensated for duplicates may now behave differently.
The v0.40.2 trade is reasonable: better llama.cpp performance and compatibility now, temporary disk overhead while backups exist, and an opt-out that is one documented command. The only real trap is a full disk on a host whose models upgrade during an unattended run. Budget the bytes, canary the behavior, and purge backups on your own schedule.