Skip to content

Infrastructure Management

Capacity planning when the cloud makes it feel unnecessary

Capacity planning used to be forced on you: hardware took months to arrive, so someone had to predict the year. Elastic infrastructure removed that deadline, and with it the habit. The planning still matters — the failure mode simply moved from running out of capacity to being surprised by a bill or by a limit nobody knew existed.

What scaling does not solve

  • Anything with a fixed size.Databases, licensed products, stateful services. Scaling these is a project with a window, and the notice you need is weeks.
  • Account and service limits.Every cloud has quotas, many are not obvious, and hitting one during a peak means an urgent support ticket during the worst hour of the year.
  • Cost.Scaling works and the invoice follows. Most cost incidents are capacity events that succeeded technically.
  • What scales badly.Adding application servers is easy; the connection-pool ceiling, the third-party API rate limit, or the single-threaded job behind them is not.
  • Anything where growth is a step, not a slope.A new customer, a campaign, a seasonal peak.

Measure headroom, not utilization

Average utilization is a poor planning input. It smooths peaks, which is precisely what capacity is for.

What to look at instead: peak utilization against the constrained resource, expressed as headroom. Not “the database runs at 40%” but “at the busiest hour we are at 70% of maximum connections, and the trend adds three points a month”. That sentence contains a date, which is what makes it a plan.

And find the actually-constrained resource rather than the most visible one. CPU is easy to graph and is frequently not the limit; connections, IOPS, memory per worker, thread pools and rate limits are less visible and more often binding.

Three horizons

Weeks

Driven by known events: a release, a campaign, a customer going live, a seasonal peak. These are scheduled by people who will tell you if asked, and being on the distribution list for marketing plans is a legitimate capacity activity.

Quarters

Driven by trend. Extrapolate the growth in your constrained resources and mark where each crosses a threshold requiring action. Anything with a long change path — a database resize, a license tier, a quota increase — needs to appear here rather than in the first horizon.

A year

Driven by business plans. If the company expects to double its customers, that is a capacity statement, and the question is which parts of the architecture do not scale linearly with it. That question rarely gets asked before the growth, and it is much cheaper before.

Find the limit before it finds you

The value of a load test is not the number it produces. It is discovering what breaks first, because that is almost never what anyone predicted.

Test to failure rather than to a target. A test that confirms you can handle twice current load tells you far less than one that reveals the connection pool gives out at 2.4×, because the second one names the thing to fix and the ceiling to watch.

Treat cost as a capacity signal

Cost per unit of work — per order, per call, per active user — is the most useful capacity metric in an elastic environment, and almost nobody tracks it. Total cost rising with growth is expected. Cost per unit rising means something is scaling worse than linearly, and that is an architectural problem showing up on the invoice before it shows up on a dashboard.

A small standing practice

  1. List the resources that actually constrain each significant system.
  2. Graph peak headroom on each, monthly, with a trend line.
  3. Mark the date each one is projected to cross its threshold.
  4. Know the lead time to change each, and start the ones with long lead times early.
  5. Keep a list of known quotas and limits, with current usage against them.
  6. Track cost per unit of work, not just total cost.

That is an hour a month and it converts capacity from a surprise into a schedule. Our infrastructure managed services produce exactly that view as part of the reporting, and our hyperconverged infrastructure work is usually where the fixed-size constraints get addressed once the trend has named them.

What can we help you achieve?

We empower your vision with innovative and effective strategies.

Aura
AI Agent

Hi there 👋

AI-powered assistant for services, careers & support

👋

Quick intro

So we can assist you better

Please enter your name
Please enter a valid email
Please enter a valid phone number
Your data is secure
Powered by AcmaCorp Solutions