Jens Klein: Plone on Kubernetes - conditions were cloudy

published Sep 24, 2026

Presentation at Plone Conference 2026

You may get some of these reports.

The site is slow. By the time you login to check, everything is calm again. You think: okay, it restarts under heavy load, call me if it happens again.

Or you see the storage is full, maybe some large videos uploaded.

Or you notice the search gets slower, every month. Adding RAM helps, but someday it does not help anymore.

"Please purge the cache." Easy enough when you have one Varnish caching server, but what if you have more? And where are they?

One cause: I could not look inside.

Having containers solved the packaging problem. But the container is still a black box.

Plone just runs stable for years. For some sites I just forget that they are there, they never cause problems. For these kinds of sites you do not need this. Plone never needed to report anything.

But now we have Kubernetes. Kubernetes is not some better way to start containers. It is a control room. It sees what you want and what is actually there, and it gets the system to your desired state.

A control loop can only control what it can measure. With Plone my conditions were cloudy. There were many unknowns. Kubernetes asked Plone questions, and Plone had to answers. Kubernetes made me answer questions I had not asked in twenty years.

Take the "it restarts under load" problem. With Kubernetes it starts a container, checks if it is ready, restarts it if it fails, and it keeps the logs, so you can look at it later.

With Kubernetes you can do rolling updates during the day, so no downtime. You can do something in Docker as well, but with Kubernetes you get it for free.

"The site is slow." You can ask: show me every request that takes longer than X seconds. You have one request. In layers. You don't guess, but you can look in details. Anyone who reads OpenTelemetry data can read my problem now. You don't need to be a Plone expert.

"Storage is full." Site with 24 years of content: 2.4 GB of text. The files people uploaded: 244 GB. You need to be able to answer: how much and where. Image scales can be on demand instead of stored. Small files you can put in Postgres. Large files, or large amount of data: you can use S3. Code: plone-pgthumbor, zodb-pgjsonb-thumborblobloader.

"Search gets slower, every month." We can ask the planner of Postgres to explain and analyze. It is just Postgres queries, so again you don't need a Plone or ZODB expert. Code: plone-pgcatalog.

"Please purge the cache." In Kubernetes you always want to have at least two of everything. You don't have a fixed URL of one caching server. I used a Helm chart a while for this, but it was not enough. But the cluster knows this. It keeps a book of what is running in the cluster. It is operational knowledge, as code. Code: cloud-vinyl.

Three of these five problems have the same reason. In ZODB everything is a Python Pickle, and this is hard to read. So store in Postgres. Code: zodb-pgjsonb, zodb-json-codec, zodb-convert.

You get scalability, fail-over, reproducible deployments. Operating Plone was Plone knowledge, but with Kubernetes not anymore. You just use what is already there.