Note · Opinion, not rated
The bottleneck moves.
Everyone is arguing about whether AI is real. The more useful question is what physically has to exist for it to keep working, who owns that, and how long they get to keep the rent.
The claim
A complex system always has exactly one thing holding it back. Relieve that constraint and you do not get a system without constraints; you get a system whose constraint has moved somewhere else. Whoever owns the current bottleneck collects an unusual share of the profit while it lasts, and the entire history of industrial capital says it does not last, because a bottleneck earning excess returns is an advertisement for capital to come and destroy it.
So the question I care about is not whether AI demand is real. It is: what is the constraint right now, what will it be next, and roughly when does the handoff happen? That is a question you can be wrong about in public, on a schedule, which is the only kind of prediction worth writing down.
Here is my answer in one line. Today the constraint is advanced packaging. By the end of this decade it is electricity. In the 2030s it stops being physical at all and becomes the question of whether the demand was ever worth the buildout.
Today: the constraint is a packaging step most people have never heard of
The intuitive story is that AI is limited by chip fabrication, and that the race is about who can print the smallest transistor. That has not been the binding constraint for a while.
An AI accelerator is not one chip. It is a large logic die sitting alongside stacks of high bandwidth memory, all bonded onto a single substrate, because no monolithic chip can deliver that memory bandwidth. The process that assembles those pieces is called CoWoS, and it is the actual chokepoint. Wafer starts are not the ceiling; qualified packaging slots are.
The numbers are worth sitting with. TSMC's CoWoS capacity went from roughly 13,000 wafer starts per month at the end of 2023 to a projected 120,000 to 130,000 by the end of 2026. That is close to a tenfold expansion in three years, an extraordinary industrial effort, and it still sold out. Reported 2026 demand runs near a million wafers against roughly 370,000 in 2024, with lead times of a year or more. The equipment needed to build new packaging lines, the thermal compression bonders, itself carries lead times of twelve to eighteen months. High bandwidth memory is allocated on a parallel track, and leading-edge logic is booked years out.
Two consequences follow, and the second is the one people miss.
First, allocation becomes the competitive weapon. Reported estimates put a single customer at around 60% of 2026 CoWoS capacity and the top three at more than 85%. When supply is fixed, the winner is not whoever designs the best chip but whoever booked the line eighteen months ago. That is a real and temporary advantage, and it accrues to scale rather than to cleverness.
Second, and more importantly: this constraint is already being solved. Tenfold capacity growth in three years is what a bottleneck looks like on its way out. Second sources are qualifying, alternative packaging approaches exist for less demanding inference workloads, and the capital is flowing exactly where you would expect it to flow. I would be surprised if packaging is still the story in 2029. Anyone building an investment case on permanent packaging scarcity is extrapolating the present.
Five years out: the constraint is electricity, and electricity does not respond to capital the way silicon does
Here is the mismatch that defines the next five years. A data center can be built in twelve to twenty-four months. A grid connection takes thirty-six to eighty-four months. Those two numbers are structurally incompatible, and no amount of money fixes the second one quickly.
The demand side is not subtle. Global data center power capacity is estimated at roughly 132 GW in 2026, up about 27% from around 104 GW in 2025, on a path toward something like 290 GW by 2030. In consumption terms, one forecast has global data center electricity going from roughly 448 TWh in 2025 to around 980 TWh by 2030; the IEA's base case lands near 945 TWh by 2030 and 1,200 TWh by 2035. US data center consumption sits around 180 TWh today with credible forecasts of 400 to 600 TWh by the end of the decade.
The supply side is where it gets physical. Large power transformers that once took twenty-four to thirty months now run four to five years, and in some accounts longer. Gas turbines are effectively sold out into 2028 and 2029, with one major manufacturer carrying an order backlog measured in tens of gigawatts. The US interconnection queue has swelled past 2,000 GW of projects waiting, with multi-year average waits and, more tellingly, historical withdrawal rates around 80%.
The mismatch, in one comparison
Time from decision to operating capacity
- Build a data center
- 12 to 24 months
- Get a grid interconnection
- 36 to 84 months
- Take delivery of a large power transformer
- 48 to 60 months
- Take delivery of a large gas turbine
- 36 to 48 months
- Bring a small modular reactor online
- Early 2030s at the earliest
You cannot compress the bottom four rows with enthusiasm. This is the whole argument for why the constraint moves from silicon to electrons, and why it stays there for years rather than quarters.
The workaround already visible is bring-your-own-power: on-site gas generation, behind the meter, bypassing the interconnection queue entirely. Some campuses are being designed to skip the grid indefinitely. That is a remarkable thing for the computing industry to be doing, and it tells you how binding the constraint is. It also relocates the bottleneck rather than removing it, since the turbine order book is now the queue.
Nuclear gets discussed more than it gets built. The hyperscaler power purchase agreements are real, but total small modular reactor capacity expected before the early 2030s is small, and the near-term nuclear activity is mostly long-dated contracts against existing plants rather than new steel in the ground. Anyone telling you SMRs solve the 2027 problem is describing a 2035 technology.
So the five-year picture: compute gets cheaper per unit and more abundant, the packaging constraint eases, and the marginal question for every new cluster becomes where the power comes from and who is permitted to burn it. Value accrues to firm generation, to the grid equipment supply chain, and to land with an existing interconnect. Sites with power become more valuable than sites with good fiber, which is an inversion of the last twenty years of data center siting.
The efficiency trap, and why cheaper compute has not reduced anyone's bill
The single most misread trend in this whole story is the collapse in inference cost.
The decline is genuinely staggering. GPT-4-class output ran about $30 per million input tokens in March 2023 and can be had for well under $0.50 from open-weight models by mid-2026. Stanford's AI Index tracked a GPT-3.5-equivalent benchmark falling from roughly $20 per million tokens in late 2022 to about $0.07 two years later. Epoch AI puts the annual rate of decline anywhere between 9x and 900x depending on which capability tier you measure. Hardware contributes its own roughly 30% annual price-performance improvement and about 40% annual energy efficiency gain.
The intuitive conclusion is that AI infrastructure demand must therefore fall. The opposite has happened, and it is not a paradox. It is Jevons. When the unit cost of something useful collapses, the set of things worth doing with it expands faster than the price falls. Goldman Sachs projects total token consumption rising roughly 24-fold between 2026 and 2030. One provider reported cost per token falling about tenfold over two years while consumption rose more than a hundredfold.
Agents are the specific mechanism. A chatbot answers a question with one pass. An agent plans, calls tools, retries, reads long context, and checks its own work, consuming perhaps a hundred times the tokens for a single task. Every efficiency gain gets immediately spent on doing more per request.
This is why I do not think falling model prices are bearish for infrastructure over the next five years. But I want to be precise about the condition, because it is exactly the kind of thing that can quietly stop being true: the argument holds only while demand elasticity exceeds the cost decline. Cost per token falling 50% a year against volume rising 10x a year means revenue grows. Cost falling 90% against volume rising 2x means it does not. That ratio is the whole ballgame, and it is observable quarterly.
Ten years out: the reckoning about whether it was worth it
By the mid-2030s the physical constraints of today should be substantially relieved. Turbines ordered now will be running. Transmission projects will have crawled through their queues. The packaging capacity being built today will look ample. The first SMRs will be operating, or they will have failed publicly, and either way we will know.
At which point the constraint stops being physical and becomes economic. The question the 2030s answer is whether the demand justified the buildout.
I want to flag an uncomfortable number here. The five largest hyperscalers are reported to be spending something in the range of $745 to $775 billion of capital expenditure in 2026 alone. That is a scale of private capital deployment with very few historical parallels, and it is being underwritten by demand forecasts rather than demand. Meanwhile the interconnection queue's roughly 80% historical withdrawal rate is a standing reminder that announced capacity and built capacity are different quantities, and the gap is usually large.
This is where the 1999 lesson bites, and it is the thing I am most careful about. Almost everyone who correctly predicted the internet still lost money, because they owned the wrong equity inside an industry they had correctly identified. The fiber they said would be needed was needed. It just arrived years before the revenue, financed by people who did not survive the wait, and the eventual profits went to whoever bought the assets afterward at a fraction of construction cost.
I think a version of that is more likely than not somewhere in this cycle. Not because the technology is fake, but because capital deployed against a forecast always overshoots, and because the assets here are long-lived and immovable. The specific form I would watch for is a stranded-asset problem: campuses built on power contracts signed at panic prices, competing against later campuses built on cheaper power, with the difference showing up as impairment.
What survives that, in my view, is whatever was scarce before the boom and remains scarce after it. Not the accelerator generation, which depreciates on a brutal schedule. Not the model, which is being commoditized in public. The land, the interconnect, the water rights, the transmission corridor, the relationship with a regulator. Boring, permanent, and hard to reproduce.
Twenty years out: what I will actually commit to
I am going to be honest about the epistemics here, because a twenty-year technology forecast is mostly a personality test. Twenty years ago the iPhone did not exist. Anyone who told you in 2006 what computing looked like in 2026 got the shape wrong even when they got the direction right.
So here is what I will and will not claim.
I will not predict the technology. Whether the architecture running in 2046 is a descendant of today's transformers, something built on different principles, or a hybrid nobody has named, I do not know, and neither does anyone selling you a view on it. Model architectures have a track record of being obsolete within a decade.
I will predict that the physical layer still charges rent. Every computing era since the mainframe has needed somewhere to put the machines, something to power them, and something to cool them. The specific silicon changed completely, four or five times. The requirement for power, land, and water did not change at all. If I have to hold one belief for twenty years, it is that the scarce physical inputs keep collecting regardless of which logical abstraction is fashionable.
I will predict that the value capture keeps migrating. It sat with hardware, then operating systems, then applications, then cloud platforms, and it is currently sitting with whoever controls the compute supply. It will move again. Betting that the current owner of the bottleneck holds it forever is the single most common way to be right about a technology and wrong about a portfolio.
That is genuinely all I am willing to say at that horizon, and I would treat anyone offering more specificity with suspicion.
What would prove me wrong
This section exists because a view with no failure condition is not a view. Each of these is observable, and if one happens I would rather you hold me to it than let me quietly rewrite the thesis.
- Packaging stays the constraint past 2029. If CoWoS-class capacity remains sold out with year-plus lead times three years from now, despite the current expansion, I have badly underestimated how hard this is to replicate, and the "bottlenecks always clear" framing is wrong on the timescale that matters.
- An architectural breakthrough cuts compute requirements by two orders of magnitude. This would invalidate the power thesis almost entirely. I consider it unlikely but not remote, and it is the single largest risk to everything above. Efficiency gains of 30 to 40% a year are already assumed; a step change is different in kind.
- Token demand growth falls below the cost decline. If consumption growth decelerates while per-token prices keep collapsing, infrastructure revenue falls even as usage rises. Watch the ratio, not either number alone.
- Power stops being scarce faster than I expect. A genuine breakthrough in generation or a collapse in demand growth would relieve the constraint early and make the 2028 to 2032 window look very different.
- The capex holds without revenue following. If hyperscaler capital expenditure keeps compounding for three more years with no corresponding revenue inflection and no writedowns, then either I am wrong about overshoot or the accounting is postponing the recognition. Either way I would want to know which.
Things I would watch quarterly, in order of information value: advanced packaging lead times, hyperscaler capital expenditure guidance and any language change around useful life, interconnection queue withdrawal rates, gas turbine order books, and the ratio of token volume growth to per-token price decline.
What this means for how I look at companies
Nothing here is a recommendation, and I want to keep the line clean. But it explains a bias you will notice in what I cover.
I am more interested in the businesses that sit underneath the excitement than the ones generating it. The scarce physical asset, the process that is genuinely hard to replicate, the toll on a road everyone has to drive. Not because those businesses are exciting, but because their advantage tends to be structural rather than narrative, and structural advantages survive the part of the cycle where sentiment does not.
The corollary is a discipline I try to hold: identifying the right industry is the easy half. The hard half is refusing to overpay for a correct thesis, which is the specific mistake that turned being right about the internet into losing money in 2001. A wonderful business bought carelessly is a mediocre investment. That is as true here as anywhere, and probably more so, because the story is so good that it does most of the work of talking you into the price.
Longer horizon. Lower confidence. Same arithmetic.
Sources
Figures above are drawn from published third-party estimates as of August 2026, and several of them are genuinely uncertain, particularly the power demand forecasts, where credible modeled ranges differ by a factor of five. Where sources conflict I have said so rather than picking the most dramatic number.
- Data center power capacity and consumption: Gartner, Brookings on IEA modeling, World Resources Institute on forecast dispersion
- Grid equipment and interconnection: Rabobank, Grid Strategies national load growth report
- Advanced packaging and memory: CoWoS capacity history, foundry allocation tracker
- Inference cost: summary of Stanford AI Index and Epoch AI data