Kimi 3: AI Capex Apocalypse or a Buying Opportunity?
The market is afraid that Kimi 3 invalidates the AI capex thesis, and spending will collapse. I don't think so.
Most of the technological breakthroughs in modern defense and aerospace industries can be traced back to World War II.
Everybody remembers the atomic bomb, but there were actually many technological leaps that happened during the war that were no less marvelous than the atomic bomb.
One of them was the German V-2 Rockets.
It was the world’s first long-range guided ballistic missile and the first human-made object to reach space. It paved the way for modern space exploration.
It wasn’t developed suddenly in the course of the war. The research trajectory was extending back approximately 20 years, including systematic military rocket development from the early 1930s. It required breakthroughs in liquid-fuel engines, turbopumps, guidance, aerodynamics, materials, and high-altitude flight.
Even after all the breakthroughs were achieved and a physical rocket program started in 1932, it took 10 years to make the first successful test in 1942. It’s estimated that Nazi Germany spent around 2 billion marks, $40 billion in contemporary dollars, to develop V-2.
10 years and 40 billion dollars.
When the Soviets occupied Germany, they found a few V-2s, though not a complete blueprint. They had to open it up, understand, reverse-engineer, and replicate.
Soviets launched the first V-2 replica in 1947, just two years after they found the original ones. Their estimated spending is around 1 to 2 billion rubles, or $3.5 to $7 billion in today’s dollars.
10 years and $40 billion on the one side, 2 years and $7 billion on the other.
This is a recurring theme throughout the history of technological advancement. Original innovators generally end up assuming excess risk and cost, while the copiers, followers, stealers- say what you want- end up with faster development and lower cost.
This is not necessarily bad. This is how the world develops collectively. It happens again and again over time.
From computers to chips, from smartphones to autonomous vehicles, we have seen this many times, and we’ll keep seeing it as long as innovation exists.
Yet, people keep getting baffled by it every time it happens. This time, it’s happening with the AI models.
Chinese model Kimi 3 was launched last week and delivered SOTA-level performance. Many people shared charts showing how little China’s hyperscalers were spending on AI compared to the US, and yet they managed to make competitive models.
This created a panic in the market, prompting investors to question the capex rationale. If the Chinese can replicate top models and even make them cheaper by spending less, how can we get a return on investment when spending shifts to these cheaper models?
There is a deep divide on this issue, and I recently observed that divide in our community as well.
Is the AI trade done? Should we sell, or is it a temporary dip we should buy?
To be honest, I was going to publish a deep dive on an attractive opportunity today, but I see that we more urgently need a direction on this AI topic before anybody could think about buying something. After all, if this is a real issue, the whole AI trade could collapse, and the market would come down with it.
I spent my early career as a competition lawyer specializing in high-technology sectors; then I made my investment career by focusing on competitive advantages in technology sectors and currently doing a PHD on the side in Competition Law & Economics focusing on technology markets. Most of my research draws on historical innovation cycles to understand how innovation shapes future competition.
Thus, I think I have a few words to say on this one.
I’ll try to explain what’s actually happening, what it means for the AI sector, and how we should think and position ourselves as investors. I think everybody will have cool heads at the end of the write-up and a grounded perspective as to what’s happening.
So, let’s start from the basics—what the actual hell is going on with the Chinese models?
How come Chinese models are so efficient?
There is no free lunch, but it’s a recurring theme in human history that we are often dragged to believe and act like there is a free lunch.
You can see this everywhere. In investing, for instance, people often think trend following would work well for them once they see a momentum in place. It doesn’t take long time to understand it’s not the way to riches. You can’t replace knowledge, research, and patience with herd following.
Most of the time, when we think there is a free lunch, it’s because the price is paid by somebody else, and this is why we don’t see it.
It’s the same in technological innovation.
Whenever someone achieves building and selling an existing product way cheaper, it’s often because the capex was frontloaded by the original innovators.
IBM spent roughly $5 billion building the System/360 mainframe. Then, Gene Amdahl, the chief architect of System/360, left IBM and founded Amdahl Corporation in 1970. His company spent just $40 million to build the Amdahl 470V/6, an alternative mainframe. Capex to discover the original know-how was assumed by IBM.
Same goes for branded drugs/generic drugs, first iPhone/copycats, Tesla/Chinese EV companies, etc. Original innovators assume the capex and go through development periods spanning years; followers/adoptors get there faster and cheaper.
Thus, whenever a company comes up with cheaper versions of the existing products and systems, we have to ask whether it’s really a free lunch, or just a lunch paid by others.
This is the question we should ask about the current Chinese AI models.
The widespread perception is that China builds cost-optimized models that are comparable to the state-of-the-art American models in capabilities:
As we said above, the perception is that they are achieving this by spending way less on R&D. For reference, the total capex of the top 20 American AI companies was way above $500 billion in 2025, while this was only around $65 billion for Chinese peers.
Yet, it looks like China has managed to catch up with the American frontier.
Kimi K3, an LLM released by the Chinese lab Moonshot.AI, is placed in the frontier level while reportedly costing way less per token compared to the American SOTA models like Claude Opus 4.8, ChatGPT 5.5, etc.
Claude Opus 4.7 costs $5.00/M input and $25.00/M output; Fable 5 costs $10/M input and $50/M per output token. ChatGPT 5.6-Sol costs $5/M input and $30/M per output token. Kimi K3 costs $3.00/M input and $15.00/M output.
When you look at the surface level, it seems like China made the best of both worlds. It looks like a free lunch.
Remember, whenever something looks like a free lunch, there is always somebody paying the price.
You have two options:
You can believe that China achieved a breakthrough unmatched by the US labs.
You can believe Chinese labs are somehow building on the paid-up innovation.
The first option is nonsensical. It’s way too audacious to think that the best of the experts in the world working in Silicon Valley with literally infinite resources are actually stupid enough to not discover doing it all at a fraction of the current costs.
That leaves the second option as the more plausible one since it’s also consistent with the history of technological development. This is exactly what these Chinese labs are doing.
They are basically distilling American models to create similarly performing models without assuming all the upfront costs related to research & development.
What the hell is distilling, you may ask.
Well, you basically take a big teacher model and use its outputs at scale to train your own model to behave similarly, thus perform similarly.
It’s basically training a smaller model on the outputs of a bigger model as the developer queries the teacher model for answers, reasoning traces, code, tool-use examples, or preference data.
Distillation can happen during post-training and in pre-training by mixing synthetic teacher-generated data into the main training corpus. The result is that we end up with a model with similar capabilities to the teacher but with substantial training efficiencies.
So, how do we know those Chinese models are distilling American models?
First, American labs are aware of the usage patterns that signal distillation. Over the past few months, Anthropic detected that several operators linked to Chinese firms carried out almost 29 million exchanges with Claude using thousands of fraudulent accounts. What are they doing? Coding Alibaba’s new front-end? I don’t think so.
Second, model specs and actual performance validate this.
Kimi K3 is a 2.8 trillion parameter model. For comparison, American SOTA’s like Claude Fable 5 and ChatGPT 5.6 are estimated to have more than 5 trillion parameters. This makes the student model perception stronger for Kimi 3.
The most important indication, however, is Kimi’s failure to pass its training efficiencies to inference.
If there was a real breakthrough, you would expect to see it in inference as well. However, despite seemingly lower cost per token, Kimi’s cost per task is at the same level as GPT 5.6:
Of course, this is not to say that Chinese didn’t do any innovation; they did many. It takes a whole different set of breakthroughs to achieve this level of performance even with distillation. They basically made the best out of their limited resources.
This way or another, these models are reality. They are here, and they look cheaper to use, though not by much. So, there are two relevant questions here:
How will this change the inner dynamics of the AI race/buildout?
What does it mean for investors in terms of positioning?
Let’s turn to these now.
Will frontier open-source kill the AI buildout?
This is the most relevant question now. After all, if you are creating a product spending billions and your competitor can somehow distill it into a comparable product at a fraction of your cost, how can you stay competitive and thrive?
First, we have to acknowledge that there wouldn’t be an open-source frontier without American capex. It’s not a free lunch. American investors are paying for it.
Up until now, they have seen positive signals related to their investments primarily through the insane ramp of frontier lab revenues. Anthropic and OpenAI’s combined ARR reportedly exceeded $70 billion as of April 2026:
This revenue ramp is seen as what justifies the whole buildout.
Hyperscalers are getting into agreements to provide capacity to those frontier labs, and they are using free cash flows to buy hardware to build that capacity for AI labs. And those AI labs are financing these capacity commitments also by investments they raise from VCs, hedge funds, wealthy individuals, and other tech firms.
So, if Chinese models can make cheaper models with similar performance, then usage can shift to Chinese models, and revenues of American labs could erode, which would send the whole AI capex down with it because of the circular nature of the deals, and AI investments will freeze. This is the catastrophe scenario the market fears.
I don’t think it’ll happen. Or more accurately, I don’t think it can happen.
First, as we saw above, the cost per intelligence of the frontier Chinese models is very close to that of American SOTAs. As long as this is the case, not too many companies in the West will prefer Chinese models due to strategic and regulatory risks. Thus, American AI lab revenues won’t stall anytime soon.
Second, there is no other option but to keep spending.
If we look at the circumstances and expect decisions to flow from the consideration of the circumstances, it misleads us to think that there are options. If we start from the actual possible realm of decisions, we get a clearer picture.
There are two possible decisions:
America cuts spending.
America doesn’t cut spending.
Let’s create a basic decision matrix considering probable outcomes of AI spending decisions of the US and China.
Now, the status quo is useless for the US. If we were talking about 12-24 months lead of American models over the Chinese ones, the status quo could be acceptable. However, in the current situation, we are looking at models basically on par, and China holds the cost advantage, however little it is.
So, obviously, from the game theory perspective, the only reasonable decision for the US is to keep spending, and even increase its expenditures, however painful it could be in the short-term.
This is not just about international dominance of the country; it’s also about the dominance of private companies. If China wins AI, Chinese companies will also win semiconductors, biotech, health-tech, cloud, etc. The only plausible decision for American companies is to keep spending as well.
If you see AI as a system and not just an LLM race, you see more clearly that it makes sense.
If you think narrowly about AI and see it just as a race to dominate the LLM market, you would have a hard time justifying all the spending. After all, you would assume constant capability improvement to justify spending, but if the capability will constantly improve, what stops people from using cheaper models that are just good enough?
Suddenly, you would find yourself in an environment where 90% of the tasks are done by “good enough” models. There’ll be many of them, and it’ll be a fragmented market. Due to this fragmentation, not enough value will accrue to the top firms to keep them making new and better models. It’ll be a dead end.
But if you see LLMs as the tip of the iceberg, that paves the way toward more general AI systems that will pervade our lives and change the world, then you can justify it.
You can assume that the system development will hit a level where it’ll be impossible to dilute all the aspects, capabilities, components, etc.
ASML’s lithography machines are an example here.
They spent €22 billion and 20 years to develop the whole system. Now, despite the blueprint being largely known for every component, no company is anywhere near replicating the whole system. ASML reportedly holds a 10-year lead over the competitors.
It makes sense to spend on system-level development to hit a milestone like this in AI.
What kind of an AI system are we talking about here?
It’s impossible to know even with our most visionary state. Think about it. Who could expect Tesla Roadster when Ford Model T was first launched? Who could expect iPhone when the first telephone was made? Who could imagine Netflix when we first came up with digital video? It’s impossible.
And honestly, it’s even more impossible for AI given the endless opportunities it opens up. We just don’t know what will be made possible by AI and how its end-state will look. It’s just a big unknown.
All this being said, it’s almost certain that the introduction of these models and the tension between the proprietary American models and the Chinese open-source models will create fluctuations in the market. I think it’s just impossible to be totally insulated from these fluctuations while still being exposed to AI to ride the upside.
If you don’t want the volatility, the option you have is cutting AI exposure altogether.
I don’t think it makes sense to completely ignore a new technology that will almost certainly have a revolutionary effect on the world. Fortunately, I think it’s possible to position yourself in a way you can hardly be a loser in the long-term, which allows you to comfortably ride short-term volatility.
I think this is the practical dimension that investors should put more thought.
So, what’s the best positioning given the trajectory of AI and LLMs? Let’s discuss.
How to optimally position in the AI trade?
We have to acknowledge two things to reason our way to the optimal positioning.
First is that demand for intelligence will always increase.
When you read this superficially, you can object and say that demand for anything never always goes up. But it does for intelligence. You think about the times AI didn’t exist and you’ll see this. Demand for intelligence has always gone up in human history. We have become more and more educated, specialized, and knowledgeable.
Because intelligence is the ultimate Jevons Paradox.
We may not know now what we can do with additional intelligence, but when we have extra intelligence, we have always found a way to make productive use of it. In old times, we were intelligent enough to create basic tools; then we were more intelligent, and we created the iPhone.
Thus, if intelligence is now something that can be represented by tokens, we can say token demand will ever increase. And given that ChatGPT was launched just in 2022, we are still in the very early innings of this ever-increasing token demand.
Second, we have to acknowledge that not all work requires the same level of intelligence.
The basic tool and iPhone example applies. Turning raw steel into a fork doesn’t take the same level of intelligence as designing the iPhone. The natural consequence of this is that we hire people from different intelligence levels to do different work.
A barber doesn’t need to be as intelligent as a chip designer at Nvidia.
Generally, people who do higher-intelligence work get paid more than those who do work that requires lower intelligence. This is how we maximize the economic benefit from a given activity. If a hairdresser got paid $1 million a year, haircuts would be too expensive for most people; the benefit of having a haircut wouldn’t be enough to justify the price.
If intelligence is now shifting to models, the same divide applies to models as well.
We will inevitably employ cheaper models for lower intelligence work and more expensive models for higher intelligence work.
The problem is that AI models develop in generations. For human beings, development starts from 0 for everybody; thus, variance is higher. Given generational leaps, variance is way lower for AI models.
Thus, at some point, even the cheapest AI models will be smart enough to do most of the work that is currently in the realm of only the so-called flagship models.
Think about Opus 4.8 today. It does almost all of the coding work that human developers could do. 12-18 months from now, non-flagship, cheaper models in the market will be performing at the same level. So, we’ll likely use them for most of the work and switch to expensive frontier models only for the top 1% of the work.
Take a look at this, and it’s obvious:
As you can see, the smaller, next-gen price/performance model Gemini Flash 3.5 performs better than the flagship models released just a year ago.
Next year, smaller and cheaper models will perform comparably or even better than today’s flagship models for most tasks. As task complexity doesn’t magically go up, spending will likely shift to these models.
We are already seeing this happen.
Chinese models have now overtaken American models in token share on OpenRouter:
If you read this as a shift from Chinese to American, you’ll be misled. All things being equal, no Western enterprise has any incentive to pick Chinese models over Americans.
If you like to see it, look at the cloud market shares. How many Western enterprises use Chinese cloud providers and how many use Amazon, Microsoft, Google, Oracle?
This is not a shift from American to Chinese; it’s a shift from expensive to cheap, and it’ll accelerate as cheaper models achieve better capabilities.
If you see it this way, you clearly see the options:
First, American labs ignore this and keep pushing the frontier harder, hoping that they’ll push so hard that the systems will reach a level not achievable by distillation.
This could work, but I don’t think it’ll stop token share gains of Chinese labs. Even if they improve only incrementally from here, they can already do most work, especially coding, in a satisfactory way. Most tokens will be Chinese, and people will turn to the frontier only when it’s needed.
In this case, what we should be aware of is that the token consumption will likely go through American cloud providers anyway. They’ll be hosting those Chinese models, and Western enterprises and individuals will be more comfortable using models through Western cloud platforms because of privacy and regulatory concerns.
The second option is that American firms take notice of this and also make cheaper models. It’ll be like distilling their own frontier by themselves and releasing cheaper-to-use models alongside the frontier.
This is easily doable for them as they are already making the frontiers from the ground up. As price-performance will be equal to Chinese models, there won’t be any incentive for anybody in the West to pick a Chinese model.
In this case, models will be distributed through Western cloud platforms as well. The only difference is that the inference margins of the cloud platforms will likely be lower, as the American labs will maximize their gross margins thanks to their long-term contracts with the cloud platforms.
So, in both cases, we’ll see token demand surge. This is not surprising, as this is what we have seen so far, and we are only in the early innings:
In the first case, however, hyperscalers will benefit both from increasing volume and higher margins. In the second case, they’ll likely have lower margins, but volume will surge anyway.
Thus, it’s actually a no-brainer to hold hyperscalers right now.
This reasoning could be passed onto the neo-clouds as well, with a reservation on their direct dealings with the American AI labs.
If American labs pick the wrong strategy and lose the token volume and revenues, their direct contracts with neo-clouds will be at risk. Hyperscalers, on the other hand, will have Chinese token demand automatically replace American AI lab commitments, so they are not exposed to the same risk.
However, a similar bullish argument can be made for neo-clouds as well. If inference demand skyrockets, even from Chinese models distributed over the cloud, we can think hyperscalers can end up absorbing excess-neocloud capacity.
However, this doesn’t change the logical order of exposure:
Buy hyperscalers; prefer higher customer diversification.
Buy neo-clouds; prefer higher exposure to hyperscaler contracts.
So, let’s look at where cloud backlogs come from, to the extent that we can know:
When it comes to neo-clouds, we know that virtually all of Nebius’ backlog is from Microsoft and Meta. Coreweave, on the other hand, has around $22 billion in commitments from OpenAI and an undisclosed amount from Anthropic in its $100 billion backlog.
The problem is that the market already reflects Coreweave’s higher risk due to both its financing structure and direct exposure to AI labs. Thus, Nebius remains overvalued, while Coreweave is discounted, and perhaps more than it deserves.
Among the hyperscalers, Microsoft is obviously the top pick at the time given lower concentration on Anthropic+OpenAI and lower valuation. I wouldn’t own Google and exclude it because of the inevitable demise of the search business. Oracle is obviously the riskiest one due to its insane OpenAI exposure.
So, assuming valuations were all fairly attractive, hyperscaler picks would be Microsoft in the first place and Amazon in the second; while Nebius would be first neo-cloud and Coreweave would be second.
Counting in the valuations changes the order. Currently, I would favor Coreweave over Nebius as the valuation discount reflects more than the incremental risk that comes from its ~30% exposure to OpenAI+Anthropic.
This is what I am already doing. I entered Microsoft two weeks ago and increased my Coreweave position. I put my money where my mouth is.
Overall, for the given reasons above, I think it’s very hard to lose money on cloud providers over the long term. Still, risk/reward could be improved by considering customers and valuations and determining priorities based on that.
Conclusion
Every major technology goes through a period where capabilities rapidly improve and then flatline. At this curve, there is a point where capabilities are good enough for most of the use cases.
Up until that point, the most important thing is the capability improvement. A substantial premium could be charged for incremental improvement in capabilities. After that point, the importance shifts to price. Willingness to pay a premium for incremental capability declines rapidly.
I think current AI models have reached that point for many tasks.
Up until this year, every new generation of models exploded the demand despite a rapid rise in the cost of intelligence. This year, however, we have seen demand shift from frontier to cheaper models, as they are often good enough for tasks at hand.
Cheaper models with frontier performance will accelerate this trend.
This is not something created by Kimi 3. It’s a frontier-level model with cost of intelligence on par with GPT-5.6. But we know for sure that there’ll be cheaper models next year and they’ll perform better than today’s frontier models.
Frontier labs should decide the strategy they’ll implement to adapt to these new token economics. This is a strategic necessity.
When it comes to AI spending, higher performance of Chinese models won’t likely lead to a collapse in AI capex. After all, without American capex, there wouldn’t be the current Chinese models, as they wouldn’t be able to find a frontier model to distill.
Cutting spending means the status quo, and China would like the current status quo better than America and American firms. So, American companies will likely see AI as a system-level development, rather than an LLM race, and increase their R&D efforts even further to reach a state that others can’t reach by just copying.
The only way out is through.
Luckily, we don’t need to speculate on this, as we know for certain that the demand for intelligence will go up, one way or another. For most of the West, this demand will flow through Western cloud providers, regardless of what model wins the token share. These cloud providers get a part of their compute capacity from neo-clouds.
So, it doesn’t take a genius to see that hyperscalers and neo-clouds with established customers offer the optimal risk/reward position within the AI trade.
If you bought speculative names that depend on the overall market vibe about AI, you may have every reason to panic. However, if you are positioned in cloud providers at attractive prices, there is no reason to panic at all.
On the opposite, I am highly inclined to see this as a buying opportunity and grow my Microsoft and Coreweave positions.
I hope this helps. Let me know your position and thoughts in the comments.

















Absolute fine piece of article .
Kimi now reportedly out of GPUs and no longer accept more workloads. I think this proves your point that hyperscaler and neoclouds are in better position than ever when cheaper model ignites explosive demands, whether Chinese or American.