Confession time: the production Nextcloud from the migration battle log has been running on K3s for a week, the repo has had an nextcloud-update.yml playbook for months, and I had never actually run it. Apps were updated the old-fashioned way, by clicking in the admin UI, or not at all. Today that changed. The run itself took 2 minutes 45 seconds. The interesting part was everything around it. Same workflow as always: I decide, my AI co-admin (Claude) checks, builds and verifies. And the very first thing it found was not a bug in the update logic, but a loaded foot-gun in line three.
First, a question I actually had to ask: if I run the main playbook, do my apps get updated? Answer: no, and that’s by design. There are two independent reasons, and the main playbook deliberately covers neither.
Reason 1: same tag, newer image
An image tag like nextcloud:34.0.4-fpm is just a name that points to an image. Official images get rebuilt under the same name whenever something underneath is patched: a Debian security fix, a PHP point release, OpenSSL. The Nextcloud version stays 34.0.4; the content doesn’t. My pods run with imagePullPolicy: IfNotPresent: if an image with that name already sits in the local containerd store, K3s never asks the registry again. So as long as the tag in the inventory doesn’t change, the main playbook applies identical manifests, K3s sees nothing to do, and the old build keeps running indefinitely. The update playbook forces the comparison with crictl pull and restarts afterwards. Same for Collabora.
Reason 2: the apps aren’t in the image
This is the bigger one. Nextcloud’s code directory, /var/www/html, is a HostPath volume on the server (/srv/nextcloud-www), and app-store apps are installed right there, next to the core. A new image brings a new core and the shipped apps. It never touches third-party apps like audioplayer, drawio or Deck. Those only move with occ app:update (or a click in the admin UI). The main playbook only enforces the app layout: these on, these off, install what’s missing. There is an opt-in switch (nextcloud_apps_update), default false.
Three layers, three levers
What changes?
Example
Lever
Nextcloud version
34.0.4 → 34.0.5, or → 35
Change the tag in the inventory, run nextcloud-k3s.yml
Rebuilt image, same tag
Security fix in PHP or Debian
nextcloud-update.yml (pull + restart)
App-store apps
audioplayer 3 → 4
nextcloud-update.yml (occ app:update --all)
Why not just fold it all into one run? Because the main playbook is supposed to be idempotent and predictable: it converges to a declared state, and running it twice changes nothing the second time. Updates are the opposite: they fetch whatever is new on the internet and swap running code. That deserves its own moment, its own backup check and its own eyes on the logs afterwards. Otherwise every small config run would quietly update apps on the side, and when something broke afterwards, you’d never know whether it was your config change or somebody else’s release.
🔫 The foot-gun: hosts: nextcloud
The play targets the inventory group nextcloud. That group currently contains the production instance, a test host, and a second instance that is in the middle of a migration: its database gets overwritten by every sync until cutover day and it is pinned to the old server’s exact version. Forget --limit once and the update rolls across all three, restarting pods on a machine that is supposed to sit perfectly still. The fix is four lines and makes the playbook refuse to start without a limit:
- name: Update | Refuse to run without --limit
ansible.builtin.assert:
that: ansible_limit is defined
fail_msg: "Run with --limit <host>, e.g. --limit cvjm - this playbook restarts Nextcloud."
run_once: true
ansible_limit is a magic variable that only exists when --limit was passed. Cheap, explicit, and it turns “oops” into an error message. First task of today’s run: “All assertions passed”. 🎉
✅ Pre-flight: four checks, two minutes
Backup fresh? The daily pull from the backup box had finished at 12:30 with DB dump and data. An update without a fresh restore point is a gamble, not maintenance.
What will actually change?occ app:update --all --showonly is the dry run. Result: exactly one app, audioplayer, jumping to 4.0.0.
Healthy baseline?occ status: no maintenance mode, no pending DB upgrade, all pods running, 190 GB free. Things you want to know before, so you can tell afterwards whether you broke something or it was already broken.
Who’s online? Five users active in the last 15 minutes. Two Nextcloud restarts and one Collabora restart mean open Office documents get disconnected. Noted, accepted, it’s a quiet Tuesday lunchtime.
⚙️ Anatomy of the run
Step
Time
Why it’s there
crictl pull Nextcloud + Collabora
51 s
Pods use imagePullPolicy: IfNotPresent, so K3s never re-pulls a tag on its own. Here the tag is pinned to an exact patch release, so this mostly confirmed that the pinned image is still the pinned image.
Restart Nextcloud
42 s
Picks up a fresh image if there is one. The deployment uses the Recreate strategy: only one pod may touch the HostPath volumes, so there is a short gap instead of a rolling handover.
Restart Collabora
20 s
Same idea for the Office server.
occ upgrade
2 s
The container entrypoint already runs it on start; this is the safety net.
occ app:update --all
2 s
audioplayer → 4.0.0. The app store only offers releases compatible with the running major, so this can’t drag in something that breaks Nextcloud 34.
Restart Nextcloud again
43 s
The interesting one, see below.
🧠 Why restart twice? OPcache and a setting that doesn’t forgive
The PHP config runs with opcache.validate_timestamps=0. That’s a performance setting: PHP compiles each file once and never checks the file on disk again for the lifetime of the process. Great for speed, but it has a nasty consequence for updates. The first restart happens beforeocc writes anything; it exists to pick up images. Then app:update overwrites the app’s PHP files on the persistent volume, but the running PHP-FPM keeps serving the old bytecode from its cache. The update looks done, app:list shows the new version number, and the old code keeps running until the next unrelated restart. For a security fix, that’s the worst kind of “done”. Hence the second restart, and only when something actually changed. Fun detail: this only bites updated apps. A freshly installed app was never in the cache, so it compiles fresh. It’s specifically updates that silently don’t take effect.
🐛 Bug #2: a task that always said “changed”
The occ upgrade task reported changed although there was nothing to upgrade. Its changed_when looked for the absence of the phrase “already at the latest version”, and Nextcloud 34 no longer prints that phrase. On a no-op run it now just prints a note pointing to the upgrade docs. Absence of a string that no longer exists is always true, so the task always claimed a change.
# before: true forever on NC 34
changed_when: "'already at the latest version' not in occ_upgrade.stdout"
# after: only a real upgrade prints this
changed_when: "'Update successful' in occ_upgrade.stdout"
Harmless today, since the app update triggered the second restart anyway. But a playbook that lies about changes trains you to ignore its output, and one day that output matters. Lesson: match on the positive signal, not on the absence of an old one.
🔍 Post-flight: trust, but verify
audioplayer 4.0.0 enabled; the list of disabled apps unchanged, so nothing got switched on or off behind my back.
nextcloud.log since the run: zero warnings, zero errors. No PHP errors in the container.
Collabora: discovery answers, the Nextcloud side of the connection is intact.
notify_push self-test: all three checks green. The push server reconnected on its own after the restarts.
5xx at the ingress: five requests, all from phones and sync clients during the seconds the pod was restarting. Expected with Recreate; the clients retry automatically.
🎯 Takeaway
The update itself was the least exciting part: one app, under three minutes, clean logs. The value was in the edges. A guard rail that makes the dangerous default impossible, a restart that exists because of one cache setting, and a status check that had been quietly lying since the last major version. Infrastructure as code doesn’t make updates magic. It makes them boring, and boring is exactly what you want from the thing that touches production. 🧰