Why teams overprovision instead of solving scaling problems
Overprovisioning is a defence mechanism against not knowing something. Here's why teams guess too big, and how real usage data replaces the guess.
"We doubled this after the outage. Has anyone checked whether we still need it?"
Someone asks it in planning, and every head turns to the same engineer. The one who sized it, or the one who inherited it. Whoever answers is also the person who gets paged if the answer is wrong, so the honest reply is usually a shrug and "let's leave it for now".
This is the conversation that Sevalla is built to change. Scaling down shouldn't come down to just one engineer's nerve. It should come down to what the service is actually using, with the platform adding the capacity back if the load returns.
Overprovisioning happens to all of us. It's a defence mechanism against not knowing something.
When we're designing the infrastructure for a system, unless we know all the moving parts and the expected traffic patterns, we can't just know what size of instances we need to create. So what do we do? We overprovision, and we monitor the situation. If we remember.
An educated guess is still a guess#
When we're sizing something, do we really know how much traffic 1 vCPU and 1 GB of memory can handle? It depends massively on the language, the framework and the runtime, but also on how the thing was built. There are so many parts we just don't know for sure, unless we built every single part by hand and load tested it along the way.
So we guess. An educated guess, but still a guess. It's usually based on past experience, with previous knowledge feeding into what we'd call an "educated guess".
The biggest unknown that pushes the size up is traffic. Unless you have comprehensive analytics running on a previous iteration, you just don't know how many requests per second you need to handle.
The guess can be wrong in either direction#
When the guess is wrong, it can be wrong either way: too big or too small. The fear is spread pretty evenly.
Too big, and the infrastructure costs more than was budgeted for. I've known people to lose their jobs because of this. Too small, and the server crashes and users can't use your service.
It's a fine line deciding what size to provision, and there's no secret formula. No cheat codes, no easy way for anyone to give you the correct answer.
Better safe than sorry wins#
The main reason teams land on too big rather than too small is the old saying: it's better to be safe than sorry. When it comes to important things, caution wins, because of a background fear of getting it wrong and the scary "scale" questions.
Sometimes it's easier to start bigger and scale back, because there's less chance of things going wrong if you overprovision.
Do we ever scale back? Sometimes#
So, do we ever scale back? Sometimes. I know that's a bit of a bad answer, but it depends.
When you look at production infrastructure, you look at the average load and check whether there's enough headroom for a traffic spike. That doesn't always happen. You might occasionally do an infrastructure review after a certain amount of work has been deployed, when you're trying to work out whether the infrastructure needs updating.
If you have a dedicated infrastructure team, they'll typically do this periodically, because they have a constant finger on the pulse of your environments. Without one, it depends on several factors. Sometimes it's a cost-cutting exercise. Other times it's just an engineer who happens to be interested in this area.
What happens on Sevalla when load rises#
When the traffic starts to climb on Sevalla, the team isn't the one watching it climb. With auto-scaling switched on, Sevalla will add an instance to the process when the CPU or memory on the running instances reaches its target, and keeps on adding them up to the maximum number you set. When the load drops away, those extra instances will go with it, and you are billed for the instances your application actually used.
The team still makes the choices that actually matter. The pod size for each process, the minimum number of instances, so that there is always a baseline running. The maximum, so a runaway job can't just scale you into a surprise bill. Whether hibernation should scale an idle app down when no requests are actually coming in. Those are settings that you choose and adjust when the numbers tell you to, not work someone has to do every time traffic moves.
It changes the scaling-down conversation massively. Doubling everything after an outage made sense when adding the capacity by hand was slow and stressful. If the platform adds instances as load rises, then the baseline can be sized for a normal day, and the busy day is covered up to a ceiling that you choose.
On AWS, or Google Cloud, or Azure, the same job means wiring up the scaling rules, the monitoring and alerts yourself, and then owning them. Your team can deploy from Git. Sevalla will handle the runtime, orchestration, networking, scaling, failover, observability, and deployment workflows behind the boundary of the platform.
Seeing what it actually used#
Checking whether you still need that doubled size starts in Analytics. Each application has charts for memory usage, CPU usage and instance count, alongside requests per minute, response times and the slowest requests. Pick a process and a time period, and you are able to see what it actually used rather than what you guessed it would be.
If CPU is sitting near 100% for long stretches of time, then the pod is too small. If memory never gets close to what the pod is providing, it is too big. The instance count chart will show when auto-scaling kicked in, and the request charts tell you whether that lined up with real traffic or with something that is running in the background. A spike every Tuesday at 2pm with a flat traffic pattern isn't a scaling problem. It's a question about what your application is doing at that time.
Over a few weeks of real traffic, the guess starts to turn into something much closer to knowledge.
When a fixed, generous size is the right call#
None of this means that overprovisioning is always wrong. Sometimes a fixed generous size is exactly the right decision. The difference is always whether you are able to explain why it is the right decision.
If you have real data from a previous iteration of the system, analytics showing your traffic and how it moves, then sizing to that data isn't a guess anymore. If the business needs a predictable infrastructure number every month, then a fixed size can matter a lot more than trimming every unused instance. Some workloads can't wait for a new instance to start when a spike arrives, so the headroom needs to already be in place. Some industries, finance for example, carry uptime commitments where redundancy on top of redundancy is a huge part of the job.
What these have in common is a reason you can use if ever asked: "We know our peak, and this covers it." That is a decision. "It's better to be safe than sorry" is still just a guess, and an uneducated one at that.
Test one service against a busy day#
Pick one stateless service the team keeps resizing. The API that got doubled after the last incident, or the worker that gets bumped every time a queue backs up. Stateless matters here, because it's the kind of service that can add and drop instances without anyone worrying about what's stored on it.
Put it on Sevalla with a sensible pod size, a minimum and maximum instance count, and auto-scaling switched on. Then send it traffic that looks like your busiest day, not your average one.
Watch three things. Does it stay responsive, judging by the response times and the slowest requests? What would it cost to run, based on how many instances it actually needed and for how long? And which jobs could the team stop doing, from the scale-up before a launch to the reminder to scale back down afterwards?
If the answers hold up, you have your evidence for the next time someone asks whether you still need it.
Next time, bring the numbers instead of the guesses.