When you build a platform, one idea is very tempting: abstract everything, so that anyone can do anything without ever touching the layer underneath. It is elegant, and once you start sharing code, it is hard to stop.
My claim in this article is that the elegance is on the wrong side. Every case you absorb makes your platform harder to operate and slower to evolve, and the cost does not grow linearly. The real elegance is a platform you can still evolve in three years.
In a previous article, Did platform engineering kill DevOps?, I argued that platform engineering leaves you with two DevOps loops living side by side, one on the platform and one on what runs on top of it. This one is about the line between them, and about what it costs to push it too far.
Abstraction is a cursor, not a switch
Take one parameter and follow it: requests and limits on Kubernetes, the memory and CPU a workload asks for and is capped at. Every application has them. The question is who sets them, and the answer depends on how far your platform goes.
- The user configures everything. You hand over a cluster with namespaces, role-based access control (RBAC) and quotas. Teams write their own manifests and pick their own numbers. Thin abstraction, real Kubernetes knowledge required, full control in exchange.
- The user picks inside a frame. You ship a chart, a template, a golden path, the supported way of doing the common thing. Sizing stays visible, framed by the defaults you chose, and perhaps by a check that rejects a manifest with no limits at all.
- The user describes an application. No infrastructure in sight, everything derived from what they declare. The smoothest of the three, and the one where sizing has left your users’ hands entirely.
Whatever position you pick, you have written an interface contract, even if nobody wrote it down. It says what the user provides, what the platform guarantees, and what nobody guarantees at all. Moving the cursor to the right buys your users a better experience, and it buys you complexity, plus the occasional late night debugging session.
Complexity does not grow linearly
Where does that complexity come from? Every parameter you take out of your users’ hands becomes a feature of your framework. It has to be designed, tested, documented, versioned and kept compatible with what your users already have in production.
The expensive part is not the feature itself, it is how it combines with everything already there. A new option can be used together with the existing ones, and those combinations are what you end up testing, documenting and supporting. Add a tenth option to a framework that has nine and you are not adding a tenth of the work you have done so far. You are adding a feature plus its interactions with nine others.
On a data platform I worked on, we generated Airflow DAGs from configuration, and every edge case was a candidate for one more option in the configuration schema. Same story with Terraform modules: each unusual need adds a parameter, each parameter is a promise you then keep across versions, and some of them end up used by a single use case. Or worse, by none.
You can flatten that curve, with strict isolation between features and a real deprecation policy. You rarely flatten it enough to make coverage free.
When run eats build
A more complex platform is riskier and slower to operate. A change that used to be local now has a wider blast radius, because more things depend on the piece you are touching. Regressions become harder to anticipate, so releases get heavier, so they get rarer.
Then the part I have actually watched happen: past a certain point, the run load eats the capacity to build. You spend so much time putting out fires in the house that nobody seriously suggests adding a second floor. The platform becomes slowest to evolve exactly when it covers the most cases. The users who asked for all those features are the first to notice, because the request that now takes two months used to take two weeks.
So the elegance is on the wrong side. A platform that covers everything and cannot be evolved any more is not a good product. The real elegance is keeping the platform simple enough that you can still evolve it in three years.
Saying no to a user
Which brings up the uncomfortable part. What do you do when a user comes with a request that only fits their case?
Saying no is hard, for two reasons. Your platform exists because its users are better off with it, so every refusal feels like it eats into your own reason to be there. And internally, a refusal is rarely a technical discussion. Your users are colleagues. The ones with enough political weight will get their feature whatever you decide. The others have nowhere to go, since your platform is usually the only one they are allowed to use.
“Yes” is the easy answer in both cases. It is also the one that, repeated often enough, gives you a platform nobody can evolve any more, which stops serving those users too. So saying no is sometimes the best thing you can do for them. And nobody thanks you for it, which is a bit sad, but that is part of the job.
The criterion I use is the one from the previous section: refuse when the case you are asked to cover adds more complexity than the usage it buys you.
Escape hatches
Saying no still leaves the user with their need, so you owe them a way out. That way out is an escape hatch.
An escape hatch lets an advanced user step out of the abstraction and implement what they need one level down: raw Terraform next to your module, a hand-written DAG next to the generated ones, a custom container image instead of the standard one. It keeps your platform simple without clipping the wings of the people who know what they are doing.
The rule that makes it work comes straight from DevOps: whatever goes through the hatch is built and run by the user who opened it. You build it, you run it. That rule has to be written in the contract. Otherwise, the first time custom code breaks at 3 a.m., guess who gets called.
Hatches are not free either. Whatever people plug into becomes a surface you maintain, and can no longer change freely. Watch how often they are used, too: if half your users step out of the golden path, the golden path is the problem.
There is also a cost you will never see. A hatch only helps users who can walk through it. Someone who cannot carry the run of custom code will scale down what they wanted to do in the first place, and a need that was given up leaves no ticket, no incident and no line in any dashboard, just resentment building up over the years.
The cursor is also a staffing question
Where you can place the cursor depends on the skills of the people in front of you. Users who know the layer underneath let you stay thin, and they will use the hatches. When they do not, you have two options: abstract more and carry the complexity that comes with it, or treat it as a change management problem and train them. And some of them do not want to be trained: infrastructure is not the job they signed up for, and no training plan fixes that. What is left then is how their team is staffed.
That position moves over time, in both directions. A team that levels up can reach one level down, a reorganization can take that ability away overnight. Hiring and training are as much of a lever as the platform itself, and neither is usually your call, so make sure the people who decide hear it from you.
So where do you stop?
Between “abstract everything” and “abstract nothing”, I see the decision as a trade-off between two principles. DRY (don’t repeat yourself) pushes you to share code, and sharing means abstracting. YAGNI (you aren’t gonna need it) reminds you that most “just in case” abstractions never earn their keep.
There is no universal answer, so what I can give you is a signal. Track two things for your platform team over time: its velocity, and the share of its time spent keeping existing things alive. If the run load grows faster than the number of users you serve, and velocity drops with it, your product is probably becoming too complex to maintain. It arrives after the fact, which is a real limitation, but it is measurable, and that beats arguing about elegance in a design review (even though we all know engineers love doing that).
Conclusion
Abstracting everything looks like the most user-friendly option, and it is, right up to the point where nobody can maintain the platform any more. Complexity does not grow with the number of features you cover, it grows with their interactions, and it lands on your team as run load that eats the time you wanted to spend building. So cover the common path, learn to refuse the rest, open escape hatches and write down that whatever goes through them is theirs to run.
When does this not apply? If your users cannot carry any run at all, escape hatches will not help them. You will either abstract more than you would like, with the cost that comes with it, or accept that some needs stay uncovered. Both are fine, as long as you choose knowingly, and as long as the conversation about skills happens somewhere.
To go further
- An article I wrote on building a cloud native data platform like a product, including DRY versus YAGNI in practice
- Did platform engineering kill DevOps?, my previous article
- There’s no such thing as a “DevOps team”, by Jez Humble