GPU & AI
Rent vs buy a GPU server: the honest economics.
The raw hourly math makes buying a GPU server look cheaper than renting it. The real math rarely agrees. Once you count the power, cooling, and colocation a GPU needs, the fact that most teams use their hardware only 30–50% of the time, and the speed at which each generation is made obsolete by the next, ownership only wins at roughly 60%+ sustained utilization over 18–24 months — with the capital, staff, and space to support it. Below that bar, renting is cheaper and carries no depreciation risk, and for many teams a hybrid of owned baseline and rented peaks beats either extreme.
In short
- Raw break-even misleads. Hourly math ignores the $50k of infrastructure a GPU needs to run.
- Utilization decides it. Buying pays off near 60%+ sustained use for 18–24 months; most teams run 30–50%.
- Depreciation is brutal. A GPU can lose half its value by ~18 months as the next generation lands.
- Hidden costs both ways. Owning adds power, cooling, and staff; renting adds egress and idle time.
- Hybrid usually wins. Own the steady floor, rent the variable ceiling.
Rent or buy a GPU server: what is the short answer?
Rent, unless you can clear a fairly high bar. Buying a GPU server only makes financial sense when you can keep it genuinely busy — roughly 60% utilization or more, around the clock, sustained for at least 18 to 24 months — and when you also have the capital, the operations staff, and the physical space to run it. That describes hyperscalers, large AI labs, and established enterprises with infrastructure teams. It does not describe most mid-market engineering organisations, and the cloud exists precisely so those teams do not have to own a fleet.
If that sounds like it leans toward renting, it does, and the rest of this piece shows the working so you can check it against your own numbers. We rent GPU servers for a living, so treat this as an interested party showing its math rather than a neutral referee — but the math is the math, and where buying or a hybrid genuinely wins, we say so plainly below.
One clarification up front, because it trips people up: "rent" here does not only mean a hyperscaler. The cheapest rental is usually a specialist GPU cloud, which can run several times less than the big three for identical hardware, so a disappointing experience renting from a hyperscaler is not evidence that owning is cheaper — often it is evidence you were renting in the wrong place. Compare ownership against a fair rental rate, not an inflated one.
The raw break-even, and why it lies
Start with the number every buy-versus-rent calculator produces. Take a roughly twenty-five-thousand-dollar H100 against a rental rate near three dollars an hour, and the raw break-even is about eight thousand hours — a little under a year of running the card constantly. Stated that way, ownership looks like an easy win for anyone with a year-long project. Most guides stop there, and they should not.
The raw math quietly assumes the GPU runs in a vacuum. It does not: a twenty-five-thousand-dollar card needs a home that can cost as much again — power, cooling, networking, space — and it needs to run near constantly to hit that break-even. Once the infrastructure is included, the real break-even on a purchased H100 server slides out to roughly eighteen months or more of continuous, near-full utilization. The honest version of the calculation is not the sticker against the hourly rate; it is the all-in cost of ownership against how busy you will actually keep the thing.
The reason this matters so much is that the raw figure is the one every vendor calculator shows, and it is the one that talks teams into purchases they regret. It is not wrong, exactly; it is just answering a simpler question than the one you are actually asking. The question is never how many hours until the card pays for itself if it runs forever — it is, given how you will really use this and everything it costs to run, whether you are better off owning or renting, and that answer is usually different.
What utilization do you actually need to justify buying?
This is the number that decides the whole question, and it is the one teams are most optimistic about. The threshold is roughly 60% sustained utilization, around the clock, held for at least 18 to 24 months. Below that, the idle hours you are paying for — in capital, power, and depreciation — make renting cheaper, because a rented GPU you are not using simply stops costing you when you release it, while an owned one keeps depreciating in the rack whether it computes or not.
The trouble is that real utilization is usually far lower than planned. Most organisations average somewhere around 30–50% across a GPU's life, because development cycles, maintenance windows, evaluation runs, and plain variation in demand leave the hardware idle a great deal of the time. The mistake is to size a buy decision on the utilization you imagine at peak rather than the average you will actually sustain — and the gap between those two numbers is exactly where ownership turns from a saving into a liability.
Owning
The hidden costs of owning
The card is only the start. The recurring costs of keeping it running add a large fraction on top of the hardware every year.
| Cost (illustrative) | Indicative monthly |
|---|---|
| Electricity (8×H100 node, ~10 kW at load) | ~$730/mo |
| Cooling | $200–500/mo |
| Colocation (per rack) | $1,500–5,000/mo |
| Maintenance, firmware, staff | varies |
| → Operating cost before hardware | ~$3,000–8,000/mo |
An eight-GPU H100 node draws on the order of ten kilowatts under load, which is roughly seven hundred dollars a month in electricity alone before cooling, colocation, and maintenance — together typically three to eight thousand dollars a month in operating cost before the hardware is even counted. Building your own facility rather than colocating pushes the entry cost into the hundreds of thousands. Across a year, these running costs commonly add something like 40–60% on top of the purchase price, which is the part the sticker never shows.
There is a staffing point hiding in that table, too. Someone has to manage firmware, drivers, cooling, and failures, and on a serious fleet that is a real role, not a side task. For a large organisation with an infrastructure team already, that cost is partly absorbed; for a smaller one, hiring or distracting engineers to babysit hardware is a cost that never appears in the purchase price but is paid every month all the same.
The hidden costs of renting
Honesty cuts both ways, so renting has its own quiet charges worth naming. The biggest on the hyperscalers is egress — the fee to move data out — which on large model checkpoints and datasets can rival the GPU cost itself; persistent storage and IP fees add more, and some providers impose minimum commitments. Then there is idle time: rent by the hour and leave a GPU running between jobs, and you pay full rate for capacity you are not using, the same waste that hurts ownership in a different shape.
Two of those are avoidable with the right provider, which is part of why specialist GPU clouds undercut hyperscalers by several times for the same hardware. Flat bandwidth removes the egress trap, and per-minute or release-on-idle billing removes the worst of the idle waste. The remaining real cost of renting is the loss of a depreciating asset you never wanted to own — and as the next section argues, that is usually a feature, not a cost.
What about depreciation and obsolescence?
This is the factor that tips many close calls decisively toward renting. GPU generations turn over every 18 to 24 months, and each one delivers a large multiple of the last generation's performance, so an owned card loses value fast as the successor arrives. A rough rule from the market is that a flagship GPU sheds around half its value by 18 months and the large majority by 30 — a card bought near twenty-five thousand dollars can be worth a small fraction of that less than three years later.
It is not only resale value; it is economic obsolescence. As newer silicon arrives, running certain workloads on the older card becomes far more expensive per unit of work than on current hardware, which means the asset you own can be outcompeted on cost by the asset someone else rents. NVIDIA holds roughly 80% of AI accelerator share against AMD's ~5–7%. The gap is less about raw silicon than about software: CUDA, NCCL, and years of framework tuning still give NVIDIA the edge for large-scale multi-node training, while AMD has become genuinely competitive on single-node inference and memory-bound models. Cloud rental sidesteps all of this by design: you always rent the current generation, and the depreciation is someone else's problem. For anyone whose edge depends on the newest hardware, that alone often settles the question.
It is worth resisting the instinct to wait for the next generation to avoid buying into depreciation, because that instinct has its own trap. The newest part is always a few months away, and a team that keeps deferring to catch it never actually ships; meanwhile the current generation handles the large majority of real workloads and its rental price keeps falling. Renting resolves the dilemma neatly — you get today's hardware today and tomorrow's tomorrow, without ever holding the bag on either.
Capex versus opex: the money you are not spending
There is a balance-sheet argument that sits underneath the cost tables. Buying a GPU fleet is a large capital expense with a long procurement lead time — often six to twelve months for current parts — that locks a great deal of money into a depreciating asset. Three hundred thousand dollars committed to hardware is three hundred thousand dollars not invested in your product, your team, or the market opportunity in front of you, and for an early-stage or fast-moving company that opportunity cost can dwarf the hourly savings.
Renting converts that capital barrier into an operating expense you pay as you use it, with no procurement wait and the freedom to change your mind as your needs change. Bursty or experimental work, short projects, or any need for GPUs this month — cloud rental is the fast path while new-silicon backlogs persist. That flexibility has real value that a pure cost-per-hour comparison misses entirely — the option to scale down, to switch hardware, or to redirect the money is worth something, and it is worth most to exactly the teams whose future is least certain.
Deciding
So when does buying actually win?
When a specific set of conditions all hold at once — read it as a spectrum of utilization and certainty, not a coin flip.
Buying wins when all of these are true together: utilization at or above roughly 60% around the clock; a stable, predictable workload for two years or more; the capital and the in-house staff to run physical infrastructure; and comfort with an 18–24 month depreciation cycle. Sustained, near-24/7 utilisation, data-residency or isolation requirements, or multi-year horizons where a dedicated or colocated node beats per-hour rental. If any one of those is missing, the case weakens quickly, and renting or a hybrid almost always comes out ahead.
The answer most teams land on: hybrid
The cleanest resolution to the whole debate is to stop treating it as a binary. Own a baseline of GPU capacity for the predictable, steady-state work — the training pipelines and inference services that genuinely run around the clock at known utilization, where ownership economics are strongest. Then rent for everything variable: peak demand, a two-week run on a new architecture, a burst for a product launch, or a day on the newest card to see whether it is worth upgrading. Owned hardware handles the floor; rented hardware handles the ceiling.
That shape optimises the cost curve in a way neither extreme can. You pay ownership rates only for the load that is steady enough to justify them, and rental rates only for the load that genuinely varies, instead of overpaying for idle owned capacity or renting your steady baseline at a premium forever. For most organisations running a mix of steady and bursty GPU work, the hybrid is not a compromise so much as the actual optimum — which is why it is where careful teams keep arriving.
Getting the split right is its own small discipline. The honest starting point is to measure, not guess: look at what your GPUs actually do over a representative few weeks, find the floor of demand that is genuinely always there, and consider owning around that line while renting everything above it. Set the floor too high and you are back to paying for idle owned capacity; too low and you rent your steady baseline at a premium. Measured carefully, the hybrid quietly beats both pure strategies for a mixed workload.
A worked example, roughly
Two quick sketches make the thresholds concrete. Imagine a team that needs a few GPUs for a three-month project — some training, a lot of iteration — and will realistically keep them busy perhaps 40% of the time. Renting is the obvious call: there is no eighteen-month horizon to amortise a purchase against, the idle 60% would be pure waste as an owned asset, and at the end of the project the cost simply stops. Buying here would mean carrying a depreciating fleet long after the work that justified it had ended.
Now imagine a team running a production inference service around the clock at a steady 70% utilization, with demand it can forecast two or three years out, an operations function already in place, and the capital to spend. That is the profile where ownership — or owning that steady baseline and renting only the spikes above it — genuinely pays, because the hardware is busy enough, for long enough, that the all-in cost of owning it undercuts years of rental. The point of the two sketches is that the same question has opposite answers depending on a single variable: how busy the hardware will really be.
Where we stand
For disclosure: we rent GPU servers, so the conclusion that most teams should rent happens to suit us. We have tried to earn your trust anyway by showing the math and naming the cases where buying or a hybrid is the better call rather than pretending ownership never makes sense. Waiting for the next generation often delays a project more than it saves — H100 and H200 handle the large majority of workloads today, and their prices keep falling. The newest parts carry long order backlogs and US export controls; we are honest about lead times rather than promising silicon we cannot source on your timeline.
What we think we do well is the renting itself: current-generation GPUs with flat bandwidth and no egress games, so the hidden costs that make some rental bills balloon are simply not there. If your numbers say own a baseline, we will help you think it through and rent you the peaks; if they say buy outright, we will tell you so even though it is not a sale for us. The honest recommendation is worth more to both of us than the convenient one.
Questions
Answered plainly
The questions teams ask before committing capital to GPUs.
Is it cheaper to rent or buy a GPU server?
For most teams, renting. The raw hourly math can favour buying, but once you add the infrastructure a GPU needs — power, cooling, colocation, staff — and the fact that most organisations run their GPUs at only 30–50% utilization, ownership only pays off at sustained high utilization over 18–24 months. Below that, renting is cheaper, and it carries no depreciation risk.
What is the break-even point for buying an H100?
On raw hourly cost alone, a roughly $25,000 H100 against ~$3/hour rental breaks even near 8,000 hours, about a year of constant use. But the GPU needs a home: with infrastructure included, real break-even shifts to around 18 or more months of near-continuous, high-utilization use — a level most teams never sustain.
What utilization justifies buying?
Roughly 60% sustained utilization, around the clock, for at least 18–24 months, and even then only if you also have the capital, the operations staff, and somewhere to put the hardware. Most workloads average well below that because of development cycles, maintenance, and variable demand, which is why renting wins for the majority.
How fast do GPUs lose value?
Quickly. GPU generations turn over every 18–24 months, each markedly faster than the last, so an owned card depreciates hard — roughly half its value gone by around 18 months and the large majority by 30. Cloud rental carries none of that risk, because you are always renting the current generation rather than holding a depreciating asset.
What is the hybrid approach?
Own a baseline of GPUs for the steady, predictable, 24/7 work where ownership economics are strongest, and rent for peaks, experiments, and access to the newest hardware. Owned capacity handles the floor; rented capacity handles the ceiling. For many teams running mixed workloads, that combination is cheaper than going all-in on either.
Do you sell GPU servers, and can I trust this advice?
We rent GPU servers, so we have an interest — and we are telling you that most teams should rent rather than buy, which happens to align here, but we will also tell you when buying or a hybrid is right and when waiting for the next generation is a mistake. Our edge is renting current hardware with flat bandwidth and no egress games, not pretending ownership never makes sense.
Work out the real number with us.
Tell us your utilization, your horizon, and your workload, and we will run the honest break-even — including when owning a baseline beats renting all of it.