The open source AI vs SaaS decision turns on two numbers: your monthly spend and your concurrency profile. Headcount is a poor proxy for both. DoiT's own engineers average $2,200 per developer per month on usage-based Claude Code, which puts a $15K–$25K node in play at roughly 7 to 12 heavy seats. But only 10–20% of a team has a request in flight at once, so ten developers are two or three concurrent sessions and half a node sitting idle.
And before you price GPUs at all, check the third option nobody benchmarks against: DeepInfra serves Llama-3.3-70B at $0.10 input / $0.32 output per million tokens. Identical weights, multi-tenant margins you can't match. The math below runs four times over, for text, images, transcription and the vector layer.
Break-even is a spend level and a concurrency profile
Two numbers decide whether you should self-host, and team size is neither of them. The first is what you spend per month on AI right now. The second is how many requests are in flight at the same moment. Everything else, model choice included, is downstream of those.
The two numbers that decide it before any model choice
A GPU node costs the same whether you saturate it or leave it idle at 3am. That's the asymmetry the whole comparison rests on: SaaS bills you per token, a node bills you per hour. So the question is whether your monthly bill has climbed past the fixed cost of a box you can keep busy.
DoiT's engineering team averages about $2,200 per developer per month on usage-based Claude Code Enterprise pricing, which puts a $15K–$25K/month self-hosting break-even at roughly 7 to 12 heavy AI seats. That's a much lower threshold than the "you need hundreds of engineers" line most TCO guides repeat. It's a spend threshold rather than a people threshold. Twelve developers who barely use an assistant don't get you there, while seven who run agents all day do.
Concurrency is the second number, and it's the one teams get wrong. DoiT's argument is that at any given moment only 10–20% of a developer team has a request in flight, so ten developers represent two or three concurrent sessions, which is half a node's capacity at best. You're paying for eight H100s and using three or four. Size by peak concurrent requests and target token throughput, then check whether that fits on one node before you price anything.
Why 'we have 50 engineers' is the wrong trigger
Headcount only correlates with spend when usage per person is uniform, and it never is. A 50-person team where 8 people run coding agents continuously and 42 use chat occasionally has the same node requirement as an 8-person team of heavy users, and roughly the same bill. Count the heavy seats.
Run these two numbers before you compare the open-weight models catalogued in our directory against anything. Pick the model after you know whether a node can stay busy, because a model that's 6 points behind the frontier is a fine trade on a saturated box and an expensive mistake on one running at a quarter load.
Open source AI vs SaaS: what self-hosting costs in 2026
Start with the rate card, because it's the only number you can look up. Everything else in a self-hosting budget is an estimate you defend in a spreadsheet.
Rented GPU rate cards: a ~7x spread on inference-capable silicon
RunPod's public pricing lists H100 SXM 80GB at $3.29/hr on Secure Cloud and $2.69/hr on Community Cloud, H200 141GB at $4.59/$3.59, A100 PCIe 80GB at $1.39/$1.19, L40S 48GB at $0.99/$0.79, and B200 180GB at $6.79/$5.98 (RunPod). That's roughly 7x between the cheapest and most expensive card that can serve inference. If your workload is a 7B or 8B model with modest context, the L40S line is the one to price against, and most TCO posts never mention it.
Lambda charges $4.29/hr for a single H100 SXM on-demand, dropping to $3.99/hr per GPU on an 8x node, with H100 PCIe at $3.29/hr and B200 SXM6 at $6.69–$6.99/hr (Lambda). An 8xH100 node at that rate is about $31.92/hr, or roughly $23,300 for a 730-hour month at 100% uptime. Nobody runs at 100% uptime, which is the point of the next number.
The commitment discount is the biggest single lever you control. Together AI charges $3.99/GPU-hr on-demand for HGX H100 and $3.19–$3.69/GPU-hr on reserved terms from 7 to 180+ days (Together AI). Call it 20% off, available only if you're confident enough in your load to sign for the term.
The provider multiplier: $12,600 to $71,800 for the same 8xH100 month
DoiT priced an identical 8xH100 node for a 730-hour month across providers: about $12,600 on Nebius spot, $22,500 on Nebius on-demand, $40,200 on AWS, $64,100 on GCP and $71,800 on Azure (DoiT). Same silicon, 5.7x spread. Your choice of provider moves the switching math more than your choice of model does, so if you're running this evaluation on hyperscaler list prices because that's where your committed spend lives, you're comparing the worst version of self-hosting against SaaS.
| Provider / card | Hourly rate | Notes |
|---|---|---|
| RunPod L40S 48GB | $0.99 / $0.79 | Secure vs Community Cloud |
| RunPod H100 SXM 80GB | $3.29 / $2.69 | Secure vs Community Cloud |
| Lambda H100 SXM (8x node) | $3.99 per GPU | ~$23,300 / 730-hr month |
| Together AI HGX H100 | $3.99 on-demand, $3.19–$3.69 reserved | 7–180+ day terms |
The line items nobody prices: idle hours, redundancy and maintenance labour
A rented node bills by the wall clock. Your API invoice tracks tokens, while your GPU invoice tracks hours you may never use. If your traffic is business-hours and weekday, you're paying for roughly 168 hours of useful capacity out of 730 and eating the rest, so multiply your effective token cost by four before comparing it to an API bill. Add a second node or a fallback API key if the service is customer-facing, because a single node has no failover. Add the engineer who patches vLLM, rotates model versions, watches KV cache pressure and gets paged at 2am. Vendors on our machine learning platforms directory will manage some of that for a margin, which is usually cheaper than the half-headcount you'd otherwise spend.
So a defensible monthly number is: node hours at your provider's real rate, divided by your real uptime, plus redundancy, plus loaded engineering time. Run all four before you compare anything.
The third option that kills most self-hosting cases
Almost every article comparing open source to SaaS puts a GPU bill next to a frontier API bill and declares a winner. That comparison skips the option most teams should pick: someone else runs the open weights for you, multi-tenant, and charges per token.
DeepInfra serves Llama-3.3-70B-Instruct-Turbo at $0.10 in / $0.32 out per million tokens, and Meta-Llama-3.1-8B at $0.02/$0.04 (DeepInfra). Same checkpoints you'd download. Same license terms. The difference is that their H100s run near capacity across thousands of customers, and yours run at whatever your traffic happens to be at 3am on a Sunday.
Identical open weights, someone else's occupancy curve
Run the arithmetic against DeepInfra's own dedicated tier and the gap gets embarrassing. A dedicated H100 80GB costs $2.20/hr there (DeepInfra), so $1,606 for a 730-hour month. At $0.32 per million output tokens, that same $1,606 buys about five billion output tokens on the serverless endpoint. Spread over the 43,800 minutes in that month, your single H100 would have to sustain roughly 115,000 output tokens per minute, every minute, just to match the API on price. A real deployment spends a large share of those minutes idle.
You buy dedicated capacity for reasons other than unit price: data residency, no third-party retention, latency floors you control, custom LoRAs, models nobody hosts. Those are real. They're procurement reasons rather than spreadsheet reasons, and you should say so out loud when you pitch it.
When renting the weights stops being cheaper
The small-model end of the market has collapsed into itself. GPT-5-nano runs $0.05/$0.40 per million tokens and GPT-5 runs $1.25/$10.00 (OpenAI), which means a proprietary model now undercuts what you'd pay to keep an 8B open model warm on your own hardware. Cheap open-weight APIs and cheap proprietary APIs are competing on the same axis, and self-hosting is losing to both.
Pick your benchmark carefully, because it decides the answer. Anthropic's spread runs from $10/$50 per million for Fable 5.1 down to $1/$5 for Haiku 4.5 (Anthropic). Benchmark self-hosting against the top of that range and it looks like a bargain. Benchmark against Haiku, or against DeepInfra's Llama, and the GPU node has to justify itself on control rather than cost.
Four modalities, four completely different break-evens
A team that self-hosts transcription almost never self-hosts text inference, and the reason is arithmetic. Each workload has its own SaaS floor price, its own GPU load pattern, and its own licensing rules. Run one break-even model across all four and you'll get a wrong answer three times.
Text inference: the widest and least predictable flip point
This is the modality where the flip point moves the most, because the API price you're comparing against spans an order of magnitude. Anthropic charges $2/$10 per million input/output tokens for Sonnet 5 and $1/$5 for Haiku 4.5 (Claude pricing), while OpenAI lists GPT-5-mini at $0.25/$2.00 and GPT-5-nano at $0.05/$0.40 (OpenAI API pricing). Pick your benchmark tier and the break-even swings from a few hundred million tokens a month to several billion.
A dedicated H100 at $2.20/hr on DeepInfra runs about $1,606 a month at full-time uptime (DeepInfra pricing). Against GPT-5-nano's blended rate, you'd need a volume most teams never reach. Against Fable 5.1 at $10/$50 per million (Claude pricing), the same node pays for itself far sooner. Decide which tier your workload needs before you model anything, and browse the models and LLMs listings if you're still shortlisting open weights.
Image generation: the licence decides it before the GPU does
Here the GPU math is the easy part and the licence is the trap. FLUX.1 [dev] weights are widely reported to ship under a non-commercial licence, which would put a paid Black Forest Labs licence between you and a commercial product, so read the terms yourself before you budget the hardware. Nobody comparing open image models to Midjourney raises it, and it's the line item that can flip the decision.
Bursty image work is what fails the test. A marketing team generating in bunches on Tuesdays leaves the card idle for the rest of the week, and idle hours bill at the same rate as busy ones. An L40S at $0.79/hr on RunPod Community Cloud (RunPod pricing) is cheap enough that steady-state batch jobs clear the bar comfortably, which is why scheduled pipelines are the shape that works here. There are 727 image generation tools in our directory, and the licence terms vary more than the output quality does.
Transcription: the modality that flips soonest
Except it usually doesn't, and the reason is AssemblyAI. OpenAI charges $0.006/minute for Whisper and gpt-4o-transcribe, and $0.003/minute for gpt-4o-mini-transcribe, which is $0.36 and $0.18 per audio hour (OpenAI API pricing). Against $0.36/hour, a $0.79/hr L40S breaks even at roughly a few hundred audio hours a month, assuming faster-than-realtime batch throughput. That's a genuinely low threshold. Then AssemblyAI prices Universal-2 at $0.15/hour and Universal-3.5 Pro at $0.21/hour (AssemblyAI pricing), and your break-even moves up by more than 2x.
The one place self-hosting still wins outright is streaming. AssemblyAI bills streaming per WebSocket session duration, including idle connection time (AssemblyAI pricing). If you hold open sockets for meetings where people aren't talking, you're paying for silence. A self-hosted endpoint charges you for GPU seconds instead of wall clock.
RAG and the vector layer runs on query volume
Vector search break-even is counted in queries per month, and the figure that gets passed around lands somewhere near 60 to 80 million queries before self-hosting beats managed pricing. Treat that as a rule of thumb rather than a rate card, since it moves with index size, replica count and how much latency budget you're willing to spend. Below that range you're paying an engineer to babysit an index that a managed service would run for less than their salary line. Embedding generation is a separate calculation again, and it's the cheapest thing to self-host of anything discussed here because the model is small and the work batches perfectly.
Four workloads, four thresholds, and no reason to move them all at once.
Buy the hardware? The 2026 memory crisis says rent
Any payback calculation you read that was written before mid-2026 is running on prices that no longer exist. The advice was reasonable at the time: buy a box, amortize it over three years, beat the rental rate by month 14. Then the memory market broke.
An ~87% price rise on an unchanged GPU
NVIDIA's RTX PRO 6000 Blackwell 96GB is the clearest example because the product didn't change. It listed at $8,565 at a retailer in early 2025 (with a $7,673 pre-order low), hit $13,250 in June 2026, and reached $16,000 by August 2026 (igor'sLAB). That's roughly 87% in about 18 months, and NVIDIA has published no official explanation.
Run that through a 2025-era spreadsheet and the payback month roughly doubles. A single-card workstation build you penciled in at $12,000 all-in is now closer to $20,000 before you've bought RAM.
Why HBM ate the DRAM supply
GPU demand is only half the story. What matters is what that demand did to the memory fabs: high-bandwidth memory now takes an estimated 23% of global DRAM wafer capacity, up from 8% in 2024, because HBM returns 3–5x the revenue per wafer that conventional DDR5 does (DatacenterDisk). Fabs moved the wafers to where the money is. Server DDR5 ECC RDIMM prices rose about 100% to 116% between Q1 2025 and Q1 2026 as a result, so the host memory around your GPUs got hit as hard as the accelerators. You feel that twice in one build: the card, and the 1TB of system RAM sitting under it.
What a depreciation schedule written in 2026 should assume
Assume the distortion outlives your purchase decision. Goldman Sachs forecast in May 2026 a 4.9% global DRAM deficit for the year, the widest shortfall in 15 years, plus a 5.1% HBM supply gap, and most analysts don't expect meaningful relief before late 2027 (Wccftech). A three-year depreciation schedule started today spends its first two years inside the shortage.
So price the buy case at August 2026 acquisition costs rather than the 2025 numbers still sitting in the ranking pages, and don't assume a cheaper refresh in year two. If the numbers only work with 2025 hardware prices, rent. Renting is how you keep the option to buy in 2028 instead.
Control and compliance: the part the spreadsheet misses
Some guarantees you can only get by running the weights yourself. Most of the ones teams cite as their reason for self-hosting, you can now buy.
What self-hosting really buys, and what it only appears to buy
Running your own inference gives you four things that no contract clause replicates: model weights that don't change under you on a vendor's deprecation schedule, no per-request data leaving your VPC, no rate limits imposed by someone else's capacity planning, and the ability to fine-tune or quantize without asking permission. If your risk register has a line about a model version being retired mid-audit, that's a real argument for owning the artifact.
The list of things self-hosting only appears to buy is longer. Data residency, encryption at rest, no-training-on-your-data commitments, SOC 2 evidence, audit logs, regional processing: all of those are purchasable as contract terms plus a region selector, and have been for a while. Paying a GPU premium for them is buying twice. The exposure driving these conversations is the penalty math, and the penalty math doesn't care who racks the hardware. Regulators have issued €7.1B across 2,245 GDPR fines through early 2026, with GDPR exposure reaching 4% of global turnover and AI Act exposure up to 7%. A self-hosted model that's misconfigured fails an audit exactly as fast as a hosted one.
The 2026 regulatory reset changed the deadlines you're planning against
Check the date on whatever compliance timeline your team is working from. The EU AI Act's high-risk obligations were pushed back by the 2026 Digital Omnibus, so the 2 August 2026 deadline that several widely-cited comparison pages still print has been superseded, and you should confirm the current dates with your own counsel before spending against them. If you approved a self-hosting budget in early 2026 to beat that date, the urgency it was built on is gone and the spend deserves a second look.
This matters most in regulated verticals. Teams shopping healthcare and life sciences AI tools are the ones most likely to have priced a private deployment against a deadline that moved.
Sovereign managed cloud is the middle path
The residency argument got weaker during 2026. Sovereign cloud regions staffed and governed inside the EU are now a managed option that didn't exist when most of the open-source-versus-SaaS comparisons you'll find were written, so check what your preferred hyperscaler offers in-region before you price a node. You can get EU-only data handling without hiring anyone to babysit vLLM at 2am.
The ordering that holds up is this. If your requirement is residency, buy sovereign managed. If it's weight permanence or unrestricted fine-tuning, self-host. If it's a line in a questionnaire that says "data must not leave our control," go read what your auditor will accept before you sign for a node.
How far behind are open weights, really?
Four months. That's the measured lag, and it's the number that should replace whatever you currently believe about open models being a generation behind.
Four months and six index points
Epoch AI put a figure on it by tracking the best open-weight release against the closed frontier on the Epoch Capabilities Index between 1 January and 28 May 2026: a 4-month lag, or 8 ECI points with a 90% confidence interval of 7 to 11. Artificial Analysis measures the same thing differently and lands in the same place. Open leaders score 52 to 54 on its Intelligence Index against GPT-5.5's 60, a 6-point gap that was 13 points twelve months earlier.
So the gap is real and it's closing. For most production workloads (classification, extraction, summarization, retrieval-augmented answers, first-draft code) a four-month-old frontier is fine. For the hard reasoning tail, it isn't, and no amount of prompt work closes 8 ECI points. Pick your workload before you pick your side of the argument. The open-weight releases tracked in our directory turn over fast enough that a model you benchmarked in March is not the one you'd deploy today.
The adoption-to-production gap that costs teams real money
Open models lose on shipping rather than on capability. The 2026 State of Open Source AI survey (N=1,410) found 79% of AI developers adopt open models but only 53% get them to production, versus 71% adoption and 63% production for closed models. They get tried more and shipped less.
That 26-point drop-off is engineering time you're paying for and not capturing in your TCO sheet. It also gets worse with scale rather than better: closed-model production rates climb from 54% at small companies to 73% at enterprises, while open stays nearly flat at 53% to 57%. The organizations with the most platform engineers are the ones abandoning open deployments most often, which is the opposite of what the self-hosting pitch predicts. Budget for the attempts that don't ship.
Migration paths: what switching actually takes
The swap itself is cheap. What costs you is the eval work you skipped before the swap, which you then pay for in production incidents.
The base-URL swap, and the documented exceptions worth naming
vLLM ships an OpenAI-compatible server, so for a plain chat completion the migration is a base URL and a model alias. Point your existing client at your endpoint, change model from gpt-5-mini to whatever you're serving, keep the rest of the code. That's the part every competitor treats as a black box and it's honestly the easy half.
The exceptions are where teams lose a week. Structured outputs and strict JSON schema enforcement behave differently across serving stacks, so anything relying on guaranteed-valid function arguments needs testing on real payloads rather than a smoke test. Parallel tool calls are a common gap. Prompt caching is provider-specific, and if your cost model assumed cached input rates on the SaaS side, that discount doesn't follow you. Long-context behavior degrades at different points than the context window advertises, and tokenizers differ too, which means your token-count-based budgets and your truncation logic both need recalibrating.
Build the eval harness before the cutover, not after
If you can't measure quality on your own traffic, you can't tell whether a four-month capability lag matters for your use case or not. Capture a few hundred real production requests with their outputs, score them however your business scores them, then run the candidate open model against the same set. Do this before you provision a GPU.
- Log 200 to 500 real requests with responses and any downstream signal you already track (thumbs, edits, retries, escalations).
- Score the incumbent SaaS output to establish your baseline. You need the number you're defending, not a vibe.
- Run the open-weight candidate through an API first, since DeepInfra serves Llama-3.3-70B at $0.10 input / $0.32 output per million tokens (DeepInfra) and costs you nothing in infrastructure to test.
- Only if quality holds and volume justifies it, move to dedicated capacity.
The hybrid default: route by request class instead of replacing wholesale
Split traffic by what each request is worth. Classification, extraction, summarization and retrieval reformulation run fine on an 8B model at $0.02/$0.04 per million tokens (DeepInfra), and the hard reasoning tail goes to Sonnet 5 at $2/$10 (Anthropic). If 80% of your calls are the cheap class, you've cut most of the bill without betting quality on a single model.
Routing also gives you a fallback path, which a wholesale replacement doesn't. Serving stacks and routers make up a good chunk of the 128 AI frameworks in our directory, and the open-source AI repos worth starting with are the ones with an OpenAI-compatible surface, because that's what keeps your rollback to one config change.
Who should switch in 2026 — and who shouldn't
Three profiles have the arithmetic on their side. Three don't, and no amount of enthusiasm from your platform team fixes that.
Three profiles where self-hosting clearly wins
High-volume transcription at steady load. If you're pushing more than a few thousand audio hours a month through a pipeline that runs on a schedule, you can keep a card busy and beat the $0.15/hour AssemblyAI floor (AssemblyAI). Batch work is the ideal shape here because you control when the queue drains, so keeping the card fed becomes a scheduling problem instead of a hope.
Regulated data that can't leave your boundary, where the requirement is contractual rather than a preference. Cost optimization isn't the goal in that case. You're buying an audit answer, and the GPU bill is its price.
Fifty-plus developers on usage-based coding assistants. At roughly 50 devs the bill lands near $110,000 a month against a node you can keep saturated, a 3–9x premium, and at ~100 devs (~$220,000/month) owned hardware pays back in about a year (DoiT).
Three where it loses on arithmetic alone
Teams under about 10 developers. At ~$22,000/month the assistant bill is a wash against a node those 10 people can't keep busy (DoiT), and a wash isn't worth an on-call rotation.
Anyone whose workload is bursty consumer traffic. Idle GPUs bill at full rate, and a spiky diurnal curve means you're paying for the peak while averaging maybe 20% of it. Serverless open-model APIs eat this profile alive.
Teams whose quality bar sits at the frontier. If your evals only pass on Fable 5.1 at $10/$50 per million tokens (Anthropic), a self-hosted open model is a different product at a different quality level, and the savings buy you a downgrade.
The 90-day test to run before you commit capital
- Weeks 1–2: instrument concurrency, not spend. Log requests in flight per minute and find your p95. If it's under 4, stop here.
- Weeks 3–6: point 10% of traffic at DeepInfra's open-weights endpoints and run your existing eval suite against it. This is a base-URL change and a model alias for most stacks.
- Weeks 7–12: rent, don't buy. An H100 at $2.20/hr dedicated (DeepInfra) gives you a real utilization number for under $1,700 a month, which is the number your CFO will ask for.
If the card sits idle more than 60% of the time at the end of that, the answer is the managed open-weights API, and you've spent one quarter learning it instead of three years depreciating it.
Frequently Asked Questions
Is self-hosting open-source AI actually cheaper than ChatGPT Business or Claude Team seats?
At list seat prices, no. Claude Team costs $20 per seat per month on annual billing ($25 monthly), with premium seats at $100/$125, and Enterprise is $20 per seat annually plus usage at API rates (Anthropic pricing). No rented GPU gets near $20 a head, so flat-rate subscriptions are the hardest thing in AI to beat with your own hardware. Self-hosting starts to compete only when your team is on usage-based billing, where DoiT's own engineers average about $2,200 per developer per month (DoiT).
At what monthly spend does self-hosting an LLM break even in 2026?
Roughly $15,000 to $25,000 a month of usage-based API spend, which at DoiT's observed $2,200 per developer per month works out to about 7 to 12 heavy seats (DoiT). The number moves more with your cloud provider than with your model: the same 8xH100 node for a 730-hour month costs about $12,600 on Nebius spot, $22,500 on Nebius on-demand, $40,200 on AWS, and $71,800 on Azure (DoiT). Headcount is the wrong unit anyway. Only 10 to 20% of a developer team has a request in flight at any moment, so ten developers are two or three concurrent sessions, at best half a node's capacity (DoiT).
How many GPU hours do I need to beat the Whisper API's $0.36 per audio hour?
Throughput settles this one, and hours only matter through the rate they bill at. OpenAI charges $0.006/minute for Whisper and gpt-4o-transcribe, which is $0.36 per audio hour, and $0.003/minute ($0.18/hour) for gpt-4o-mini-transcribe (OpenAI). On a RunPod L40S 48GB at $0.79/hr Community Cloud (RunPod), you need better than about 2.2x real-time throughput to beat $0.36, about 4.4x to beat the mini tier, and about 5.3x to beat AssemblyAI's $0.15/hour Universal-2 rate (AssemblyAI). Idle GPU time counts against you, so the deployment only wins with a steady queue rather than bursty daytime traffic.
Is open-source AI free to use commercially?
No, and the image models are where teams get caught. FLUX.1 [dev] weights are reported to carry a non-commercial licence, which would put a paid Black Forest Labs licence between you and a commercial deployment, so read the terms before you plan around them. Even with fully permissive weights, the compute isn't free: DeepInfra serves Llama-3.3-70B-Instruct-Turbo at $0.10 input and $0.32 output per million tokens, and Meta-Llama-3.1-8B at $0.02/$0.04 (DeepInfra), which is cheaper than most teams can run the identical weights themselves on a $2.20/hr dedicated H100.
Do I have to self-host to be GDPR and EU AI Act compliant?
No. Compliance is about data location, processing agreements, and documentation, none of which require you to own the inference stack. The stakes are real (regulators have issued €7.1 billion across 2,245 GDPR fines through early 2026, and exposure runs to 4% of global turnover under GDPR plus up to 7% under the AI Act, per Kiteworks), but EU-region managed and sovereign hosting options now cover most of what a data protection officer asks for. Self-host when a contract or regulator specifically forbids third-party processing, not as a general hedge.
Should I buy GPUs or rent them during the 2026 memory shortage?
Rent, unless you have a multi-year workload that keeps the cards busy. NVIDIA's RTX PRO 6000 Blackwell 96GB went from an $8,565 retailer listing in early 2025 to $13,250 in June 2026 and $16,000 by August 2026, about 87% more for an unchanged product (igor'sLAB). The cause is memory: HBM now takes an estimated 23% of global DRAM wafer capacity versus 8% in 2024, and server DDR5 ECC RDIMM prices rose roughly 100 to 116% between Q1 2025 and Q1 2026 (DataCenterDisk). Goldman Sachs forecast a 4.9% DRAM deficit for 2026, the widest in 15 years, with relief unlikely before late 2027 (Wccftech), so plan on elevated purchase prices and take the rental market's spread instead, from $2.20/hr for a dedicated H100 on DeepInfra up to $4.29/hr on Lambda.







