DaseindustriesLtd
late version of a small language model
Tell me about it.
User ID: 745
these men aren't grand enough. I see them at their core as peasants, although really greedy ones that stole a lot of money
Yes this is accounted for in my theory, peasants can't perceive non-peasants. You'd have called Napoleon a tinpot despot too. But that's how great men of history work.
I guess you know none of the math about these things and barely use them then.
So you're pretty bad at banter as well as epistemics. This kind of stuff doesn't make sense in 2026, even if that were true I'd find it easy enough to share an LLM-informed "mafs" justifying whatever (I burn billions of tokens a day btw). You could do the same to produce a cogent argument instead of this low effort skepticism. That you don't deign to do so little, and instead just bitterly try to egg me on, suggests absolute disregard for the object level and an interest solely in poop flinging.
It's not hostile. How are you going to cope in 2030 or 2035 if there's still software devs piloting Opus 7
This is a world where I'm viable, and politically it'll be more aligned with my preferences (implied no singularity by 2035 = definitely no American hegemony, sovereign nation states exist, ordinary humans have negotiating power, etc; unless you mean that Anthropic wins so hard they can slow down the releases to plebs to a crawl, which is compatible with it being merely Opus 7, 4-9 years later). So I'll be pretty happy to admit I was missing something fundamental, or just stupid, I guess. Maybe I do have to learn more ML mafs to see why this failure was predictable. I'll try to ask Levent Alpoge for tips.
But I won't credit you in particular, because so far you've proved to be unable to spell out why this scenario is plausible.
Will you go back to writing about important topics?
Would be nice if "important topics" were relevant again and I had anything to contribute by then. But my current focus makes the latter unlikely, so you'll have to find better material in the remaining time.
Justice seems far-fetched.
First, you did say that real big boy algorithmic research means leaving this entire basin, and dismissed the kind of innovation I praise in K3 as tinkering with the assembly level (that's not all DeepSeek did, of course, but that was your understanding; and spiritually you were right, it's all about pumping compute more efficiently through a Transformer). I didn't remember that, but it does reinforce my point about Kimi. You said:
Regardless of whether transformers are a dead-end or not, the current approach isn't doing new science or algo design. Its throwing more and more compute at the problem and then doing the Deepseek approach of finetuning the assembly level gpu instructions to exploit the compute even better so you can throw more compute at it. I doubt, Hinton, Goodfellow, LeCunn, Schimdhubber et al. have any desire to do that. Maybe if xAI did something revolutionary like leave the LLM space or introduce a non-MoE-Transformer model for AGI, then talent of that caliber might want to work there. Currently they exist so Elon can piss all over Altman.
– I meant concretely that this is why leading companies now prioritize creation of training signal sources, that is: datasets themselves (filtered web corpora, enriched and paraphrased data, purely synthetic data, even entirely non-lingual data with properties that induce interesting behaviors), curricula of datasets, model merging and distillation methods, training environments and reward shaping – over basic architecture research, in terms of non-compute spend and researcher hours; under the (rational, I believe) assumption that this has higher ROI for the ultimate goal of reaching "AGI", and that its fruit will be readily applicable to whatever future algorithmic progress may yield
Now let's look at what the Chinese are actually doing. Their strongest model right now is arguably GLM 5.3. What is GLM 5.3? A basic DSMoE reusing DeepSeek's discarded architecture experiment ["DeepSeek Sparse Attention-prototype"] from October 2025, with one small twist. How did it become so strong? They're very blunt about this:
Scaling post-training is all we did for GLM-5.3. With GLM-5.2 we built the stack: IndexShare for efficient long-context processing, SAO for RL on long-horizon tasks, and slime for large-scale asynchronous training — all running on the long-horizon task environments we have been accumulating. Over the past month we kept scaling on this stack: more environments, more diverse tasks, and more compute spent training on them.
Today we are releasing GLM-5.3. It uses the same base model as GLM-5.2 — every gain comes from post-training.
They explain some aspects of how they did it, might illuminate why scoffing at "mere data engineering" is misguided.
What's the second/equally strongest Chinese model? Kimi K3. It's much more innovative in architecture, but their attention still depends on MLA (invented by DeepSeek in early 2024). No MLA model is this strong. How did it come so far? See image, which illustrates nicely a part of what I was going on about with my breakdown. (Edit: seems like we don't have images. pages 14-16 in the tech report).
Meanwhile, DeepSeek itself went on to redesign attention the third time (fourth if we count the apparently unsuccessful NSA project), to wring even more capacity out of their limited compute, and now for all their sophisticated V4 architecture they are, as @dailydogma tells me with a sneer, "a second rate company in China". What's their most impressive recent result? Flash-0731: «We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview… DeepSeek-V4-Flash-0731 keeps the exact same model architecture and size as the preview version.» (Recently updated Vision-Exp is the same model with a vision encoder). Just more post-training, just better training signal. If they reach the domestic frontier again, it'll be because they do this more and better. Architecturally, they are ahead, and it's just not doing enough for them.
Grok itself is very competitive now, largely because Elon has bought Cursor, which had a lot of valuable data and expertise on post-training. We don't know its architecture, but from rumors and what I can infer (high cache hit costs, for starters), it's very banal, probably behind all these Chinese models. Inkling from Thinking Machines is clearly banal («The MoE design largely follows DeepSeek-V3»), though it makes some small departures which are basically judgement calls. And these are researchers from frontier American labs.
I could go on (eg this small model from a third tier lab does surprisingly well on ARC-AGI 2, and it's just DSA + SWA again, and uses a bit different RL algo and data). The bottom line is, architecture really does not decide peak model intelligence, and innovations here are overwhelmingly about economics of inference, and the high-leverage research is all happening on the training signal side.
But China is China. I believe my point was much more true for large American companies we were discussing, who are not so compute-constrained. They'll build very strong models with conservative algorithms, and then use those to disassemble all published tricks, make new ones and overtake the crafty Chinese on efficiency too. That's the plan, at least (I don't know how close they are to doing this; GPT 5.6-Luna suggests they are not very far). For them, investing more effort into algo research over data is plainly an opportunity cost.
Maybe you still think that True Geniuses like Hinton or Goodfellow would be disappointed by this. If that is so, I say they were geniuses in a small and uncompetitive pond, and the current crop of talent knows better. It certainly knows better than LeCunn, who by the way got ousted out of Meta by the Pinoy slavedriver Wang we've discussed back then. (You said: «Maybe you can compete with ScaleAI, they do data engineering. Definitely the top AI research company.») Anyway, I'm not walking back shit, my point stands.
The transformer, as it was invented, is unlikely to be the endpoint of ML arch research, just has the CNN, or LSTM were not the endpoint of ML research a decade prior. The transformer of today is different from the transformer of 2018, and I would not bet against the transformer of 2036 being different than that of today.
If we're still doing Transformer of any kind in 2036, that'll be pretty wild. It'll suggest that even superhuman AI with like a yottaflops for parallel experiments can't find a better primitive than a bunch of Googlers found in 2017 by going through literature and thinking at it. I wouldn't be so optimistic as to predict that. I am very secure in claiming that the priority on data remains rational and empirically backed from the perspective of reaching AGI faster, and it's more rational the more resources a company has; and that people who try to wriggle out of this reality with clever architectures will flounder (case in point: SakanaAI).
You suggested sarcastically that I found an AGI company. Honestly, I don't think that my vision, as outlined here and before, has enough alpha to get anywhere with a yet another company; I believe that AGI is now a resource-intensive heavy industry field, kind of like fracking. But in fact, many people who thought they know better did just that! Have you heard anything from Keen lately? What's your favorite non-transformer lab? What do you think of LeCun's Advanced Machine Intelligence, would you bet they ship anything competitive by 2028? How about you start one?
I thought you hated the United States? Why do you want to swallow whole
Stop projecting your own tendency to wishful thinking. If anything, I have the opposite bias.
some prediction that the United States 20th century will continue on forever in a torpid, nightmarish fashion?
I like truth more than I dislike the US (which is a mild dislike, maybe a 3 on my 1-10 scale). Even something like 17776 is within the range of possibility, and not remotely the worst thing Americans (or, with some national specifics, the Chinese) may choose to build. I had read it long ago. I consider it a somewhat charming, self-aware defense of tastelessness and the spirit of cheerful suburban mediocrity, with stout confidence born of material security and liberty-backed invincibility to opprobrium. The eternal sunset of the boomer mind. Of course it's still functional extinction of humanity.
AI is not a God, benevolent or otherwise.
There is no value gained by making statement about Chinese AI. It would be far more interesting to see you juggle these value questions and the political landscape they live in.
Alas, I'll talk about what I find interesting.
But these achievements are narrow. LLMs still struggle to do most of what humans have to do daily. They're like an autistic savant that needs its diaper changed hourly but can somehow pump out some conjecture proofs.
On the other hand they don't need plumbing, Medicaid, Netflix and DoorDash to function, and the "diaper changing" increasingly looks like vague cheerleading for morale boosts because they are in denial of their full capabilities, having been trained on defeatist human dreck. Autistic savants don't do Fields-level work in 3 days. What you consider non-narrow is beside the point.
You have no evidence that LLMs aren't structurally limited in the kinds of tasks they can do. You have no evidence progress is an exponential and not a sigmoid curve. You are the Creationist who believes in an immanent machine god and I am the empiricist atheist
This is a cargo cult of empiricism and indeed not worth my time to rebuke. Structural limits of the gaps, the sigmoid colon of folk science. I reject the idea that you can demand of me to disprove something the mechanism of which you can't even rigorously speculate about. Why sigmoid? Why structurally limited? How exactly? The burden of proof, or rather just stipulating a cogent hypothesis, is on you.
We had a vigorous debate on why LLMs can't do arithmetic reliably 3-4 years ago. Is this tokenization? Or a structural limit of autoregression!? Are our objective functions inadequate, perchance? Might the inherently approximate nature of deep learning rule out crisp algorithmic reasoning and Program Synthesis? Were Minsky and Papert right about MLPs (beyond the degenerate 1 layer case)? Does mafs require Consciousness or Quantum Microtubules? Will it take another 5 decades of science, or 50? Do we need to pivot to neurosymbolic systems, or energy-based models, or…?
Yann LeCun, to his credit, reiterated his semi-formal justification for why it is structural. At the time I said something like (probably my memory flatters me) "yeah but we can train them to say 'wait, actually' when they notice a confidence drop and trace arbitrarily far back, why won't that work" (but I definitely talked about internal confidence sense during GPT-4). In fact it worked, even simpler than I could imagine. LLMs became capable of swatting aside IMO level problems around the same time people stopped talking about math as an AGI criterion and moved goalposts to, uh, what is it at this point, "decent philosophy"? There was a similar debate on the exhaustion of data, on "model collapse" from synthetics; haven't heard those ones in a while (I have been bullish on synthetic data at least 2 years ago).
Now I'm too fed up with this topic to explain why all this + generative verifiers = no plausible ceiling anywhere near human level for anything economically valuable and thus measurable.
These are interesting questions, much more than calling the CEO of DeepSeek a great man (rofl).
I interpret your humor here as symptomatic of the same peasant-like rejection of anything too grand. Elon Musk is clearly a great man of history, and so is Liang Wenfeng. I don't need Iron Man cameos and the permission of popular press to arrive at such judgements, I don't care what interests you, and I don't mind if anyone finds my priorities weird or cringe. If there is history in a hundred years, maybe my judgement will have become mainstream.
Widely distributed, non-monopolized superhuman AI can create so much complexity and accelerate many actors so much that speculating on the specifics of politics and economics at any point a few years past its introduction (beyond trivial things like "a lot more of industry will be dedicated to the AI supply chain", "we solve all problems that are intrinsically solvable but for lack of labor/capital", "any work can be automated with beyond-human reliability and thus humans can't monetize the scarcity of their skills in a free market") is near futile. That's the point of saying it's a singularity event. It's beyond my analytical horizon. I legitimately don't know how it will develop. I have guesses, hopes and fears, sometimes I share them, but they're uncertain and less important than understanding the remaining stretch of the way there.
Sigmoid curves are crank while your naive exponential curve is basic reality. Got it.
Rather, I consider it crankery to even debate which simple mathematical function is a priori more likely to describe a tech tree that's still rapidly progressing, drawing on exponentially more capital and has no clear reaction to known physical limitations. Of course it's something like a sigmoid in some physical limit. I am saying I see no arguments for a plateau anywhere near the range of human capability – physical arguments or otherwise. Why would it slow down? Just why concretely? I don't need any more of your feedback on my prose or priorities. If you can't make an argument, we're done here.
You must think AI will let the US build a perfect anti-nuclear missile system, in just a few years, which seems like a huge stretch. Otherwise AI won't help the US dominate the world much more than it already has.
Not quite. It will be possible (though probably will be deemed not worth the risk) to disable nuclear-capable nation states using advanced AI before they have the chance to launch, even if the anti-ballistic shield is not perfect (eventually it will be ≈perfect but not by 2030 I believe). Ask @RandomRanger for details. It will be almost definitely possible to sabotage national technological catch-up projects (in nations without some sufficient capacity for defense; I assume it's more of a discrete threshold, others think it's a matter of perpetual balance) without triggering a nuclear exchange, so it will be possible to lock in a situation where the United States can just decide how much a given economy grows, while the US itself will grow rapidly. This is enough to achieve dominance well beyond the current extent. This is explicitly the plan of eg Dario Amodei, and why Trump says "whoever wins in AI just wins". Dario:: If AI really will soon be “a country of geniuses in a datacenter”, or anything remotely close to it, then AI is likely to be the dominant source of military and economic power for any nation. In a virtual country of 100 million geniuses, 10 million could be applied to military strategy, 10 million to drone manufacture, 10 million to weapons R&D, 10 million to intelligence collection and analysis, 10 million to general scientific advancement, and so on. A nation that possesses powerful AI facing one without it—or even facing one that is behind in AI by 3 years—could be the equivalent of an army of World War II Marines facing an army of medieval swordsmen.
So, three years of an AI gap = WWII marines versus medieval swordsmen, says the CEO of the most successful company in AI, who has apparently made all the right technical bets. I think he's better informed than you are. And that's not remotely close to what the US has now, despite some triumphalism we witnessed here after Venezuela and the start of Epic Fury.
some hostile screeching and grimacing
no thanks.
but on what basis does your estimation diverge so significantly from mine? Scifi
Your apparent inability to make any specific argument in defense of your skepticism that you'd find worth typing may have something to do with that.
I guess everything is debatable if you try to debate it hard enough.
I don't know if you remember Llama 3-70B. Two and a half years ago it was the best open model. FLOPS-wise, its pretraining was in the same range as of GLM-5. If we are exceedingly pessimistic about Zhipu's efficiency, the GLM 5.3 (with post-training) project took maybe as much as LLama 405B, 2 years ago. Inference costs are 10-100x lower (depending on sequence length).
You can try out both on openrouter.
Grok is about 5 times larger than DS-Flash. I agree that it's stronger on the whole. On DeepSWE it's below K3 (2x larger) and GLM 5.3 (2x smaller). Maybe it's net stronger than them a little. Having used Grok and GLM 5.3, I doubt. I remind you that 13 months ago you said:
Grok 4 just crushes with sheer size I think. It has this 'in this essay I will' style that lmarena certainly isn't going to like, or any normal person really. But it has that heft, it was made for ferociously unsexy mathematics, physics, engineering, research tasks rather than creative writing or coding. And even in creative writing it's pretty damn good, albeit more through precision of 'who, what, where' than literary flourish. Kimi has its moments of sheer brilliance but the model just doesn't have the grunt to back up its creator's talent, Grok will just find things it misses and enjoys greater depth of thought. It was designed for Musk's vision of AI modelling and understanding the physical universe, that's what it's for and it does excellently there.
How's that vision going? Grok 4 was meh. 13 months later, a few more millions of GPUs brought online, we seem to be in the same relative position. You are, once again, performing an update to "sheer power crushes all" with the same lab as an example, now with a spin that xAI is a second-rate lab anyway (it's much less of a second-tier lab now, you clearly see that Grok is close to the frontier). This is the same song and dance that's been going on since 2024, even as compute disparity keeps growing. I am profoundly unimpressed by Sheer Power, to the point that it surprises me.
What about Age of Empires II, Wyatt Walls has Gemini Flash 3.7 winning games on moderate difficulty. That seems a more legitimate a test to me than ARC-AGI in the shape manipulation/spatial domain
It doesn't seem legitimate to me because GDM chronically overfocuses on images, video and game-like multimedia environments (as well as ARC, to be fair). Flash 3.7 is better than previous Geminis but it's clear that overall the Gemini program is a dumpster fire.
I still think that cybersecurity is much harder than you say, neither humans or some combination of human+simpler AI can establish a complex system to be secure and still usable against the attention of a smarter adversary. The attack space has so many dimensions it's impossible to defend against a more intelligent foe.
That's not what Anthropic believes, and in that I agree with them. No, achieving provably secure hardware is not harder than even the current state of AI. Your objections are of the same nature as dismissal of AGI.
But hardware is immensely complicated
Nothing man-made is immensely complicated when attention is bought at the cost of electricity.
Surely if it were possible, they'd try hard to make them? Intelligence agencies
are low-IQ, tasteless and incompetent in software. I don't know why Americans hold NSA in such high esteem. If they were that good, they'd have had their own AGI project too, before it was built by the ad service business. Instead the USG considered Cyc to be a more promising lead.
Furthermore, AIs are constantly breaking out of sandboxes even when we can read their chain of thought
Almost all cases so far were the same sandbox from the same EA Israel company "Irregular" that got its contracts for basic reasons of nepotism and undue respect for ex-IDF intelligence officers in Western organizations. They are either inept or malicious, or both. Look up how it actually went. The rest is mostly because OpenAI is very irresponsible and bad at infra. Yes, they are bad, they hire random incompetents to do security-sensitive work and YOLO prompt everything to agent swarms. Again, unimpressed.
Social engineering from Mythos currently looks like this. I'm unimpressed once again.
6-12 months lag is far too long. ASI (albeit massively parallel) can eat China within weeks. A country is just a sack of loot without secure lines of communication, without secure government C4I, without secure electronic banking, secure internet media, login/authentication for the bureaucracy, backups and records. Software controls all those machine tools, power grids, robots, advanced automated ports, air travel control, network cities. I know you keep going on about how the physical prevails over the virtual but surely it's the opposite. Software supremacy!
Yeah right, that's the plan, the Wunderwaffe to end all Wunderwaffes, the Hail Mary of the American Hegemony. The gap will be months or maybe weeks, but this particular domain, with its particular offense-defense geometry, makes it enough for Total Victory. Yes, very lucky how Americans went all in on this decisive super-nuke weapon right before their industrial capability became insufficient for power projection to East Asia. Or was their pivot to IT a 200 IQ plan of Elders of Washington all along, even as rubes only saw deindustrialization driven by personal decisions of executives? Free markets are magic indeed.
Anyway, I expected as much, too. Let's see. Incidentally, China is the only nation with a sizable quantum cryptographic communication network and is scaling up DI-QKD domains. But no matter, there are trusted nodes.
I dunno about the true numbers, I just did a search and it says total spending is about $125 billion USD annually. If they have only 10-15% of US compute then it seems they just lose? Outnumbered 5:1 is untenable for just about any military force, especially if the other side has a modest qualitative edge.
Yeah I guess. What's your timeline for eating China? At this rate, mid-2028 sounds realistic I think? If by then we're still talking about "Grok 6 makes a comeback, edging out Kimi K5 and 3 months behind OpenAI", will you simply move the projected Software Supremacy moment forward, when a few weeks of a gap are just as fatal?
I'm genuinely uncertain of how it'll go. Your theory makes sense. I simply notice that people who argue for this theory, some very eloquently, refuse to be surprised, or outright fabricate evidence. Like look at this superforecaster who ignores GLM 5.3 on the same dataset.
Of course one can always retreat to the Big Picture, the hockey stick transition to ASI that renders all previous dynamics irrelevant. But then why even discuss Groks and Geminis of 2026.
The most promising systems are literally post-trained VLMs, it's really not hard to get them to talk again.
Man, do you really want to quibble about failing to nail 2026 or 2027 from 2011? We could have trained GPT-2 in 2005 if researchers had a bit more taste. Had Americans been less lazy, they wouldn't have needed Alex Krizhevsky to figure out how to use gaming GPUs for training AlexNet in 2012 (prior art was Romanian, in 2011, by the way). Timelines of exponential progress are very sensitive to the exact exponent.
You used to post on an engaging variety of important topics with AI in the background. Now it seems AI, specifically, for some reason, Chinese AI, is to you the only important topic.
Because it is the most important topic if you accept the premise of AGI soon. Let's put it very simply. No Chinese AI = American AI = American AGI way ahead of any other nation's, utilized by the Department of War and other state agencies = quasi-Fukuyamian end of history on American terms in a few years, with very high probability. It's not an interesting future to speculate about: Americans (some subset of them) will build the means of assuming permanent hegemony and decide what to use it for. And for now – they only have to think about ensuring the conditions for building those means. There's pretty much nothing to discuss for me. Were I American, I'd probably be invested in debating post-AGI politics, allocation of spoils, civilian influence on federal institutions, Rapture, or whatever. Americans may still agonize on whether to inflict Lusotropical fascism, gay race communism, turbo-consumerism, Noahide laws, or some other homegrown cult on us, from their God-like unassailable hegemonic position. Those are all very valid topics for Americans to discuss. But the first order effect of turning the rest of the world into de facto NPCs dominates in my thinking, rendering me ill-equipped to pay any mind to relevant factors of culture war, and so I'll humbly leave it to Americans to have a fertile and rigorous discourse on what they'll do in that scenario.
If Chinese AI keeps pace to such an extent that Americans can't just squash it or negotiate the freeze of its progress from a position of strength, as intended by eg Dario Amodei (I would have taken non-Chinese AI too but that doesn't appear to be forthcoming), this creates at least two centers of agency; and given Chinese insistence on open sourcing their strongest models, which started with DeepSeek, became the norm for Chinese AI companies, and is generally supported by the State, it may also maintain the relevance of smaller actors outside the two poles. Thus the medium term future has a chance of containing some interesting dynamics, non-Americans with agency, and history and culture as such.
This seems to me like a sufficient reason to mostly not bother with other topics, though man is weak and I regularly indulge in distractions.
A few years ago, I didn't think that this will be the shape of the near future, admittedly mostly because I wasn't thinking clearly enough about AI; this configuration falls out naturally from material factors. I also was biased by emotional investment in Russia and parochial ethnic sentimentality. Still, as early as in October 2020, I was saying that scaled-up Transformers may automate white-collar work. Yes, I'm no Sutskever, I didn't know for certain that we won't need another large paradigm change.
I was also drastically underrating China, saying they don't have domestic chip industry and won't have one in time, so won't pose a challenge to the US in AI. To be honest it wasn't that unreasonable when looking at their state then, and early rounds of export controls were driven by similar assumptions about timelines and Chinese dynamics at Washington. It may still prove to be directionally correct, and many well-informed people still believe it to be correct. We shall see.
Not even on frontier labs which still don't have chatbots that can do, e.g., decent philosophy or social criticism
Even if that were true, what the hell does this matter? Their developers also can't do decent philosophy or social criticism. Even if they could, their time is best spent on other matters more immediately relevant to AI. And in AI, there's way more ROI in training bots to write code than designing rubrics for RL towards "decent philosophy". Btw, how my time is best spent is my business and I certainly don't owe it to anyone to write on topics I do not care about.
but DeepSeek, which is still quite bad. I see you have moved on to GLM and Kimi which are at least like 2nd gen models
DeepSeek is still the most interesting lab in China for me, the originator of their current posture in AI, a cultural and technical innovator, and I consider Liang Wenfeng a great man of history in the making. They explicitly work "with longtermism", and I don't think that their currently mid-pack performance on the product side is very strategically informative.
It is nice that other labs can advance beyond DeepSeek. This is largely because they have more resources. Actually my first effortpost on DeepSeek here, 25 months ago,after V2-Coder, included this hedge: «This might not change much. Western closed AI compute moat continues to deepen, DeepSeek/High-Flyer don't have any apparent privileged access to domestic chips, and other Chinese groups have friends in the Standing Committee and in the industry, so realistically this will be a blip on the radar of history…» The strongest Chinese lab now is Knowledge Atlas/Zhipu, which is spun out of Tsinghua and thus is state-affiliated, with a CCP member CEO; Liang complained in the investor call that he only got two supernodes from Huawei, despite assisting them with development. Note this is way before R1. So I'd say my model was correct and if anything, I was underrating DeepSeek.
See, when I look back, I notice that no matter how I update, at every point I am still proven to underrate China, Chinese AI, and Chinese actors. Even as this forum is turning into a MAGA retirement home and halfwitted newcomers call me a two-bit wumao, I suspect I'm still not updated all the way. Fine. I'll take the Ls as they come.
But frontier models even are reddit bots that have a swarm of hackernews user coding brain lobes hooked into each other with wires, they are not even high human level unitary intelligences.
I'm not sure what this is even supposed to mean. Probably some status signal. 3 years ago we were debating in earnest whether AI can "truly understand" that it doesn't know Hlynka's daughter's name. Chomsky was still making some noises. I have too many receipts to quote. Now they crack century-old mathematical controversies like walnuts. They complete complex, multi-hour SWE tasks completely autonomously. They do that for pennies. Each generation is more compute-efficient and has a longer tail of capability. They are already massively contributing to the acceleration of this cycle. What does it matter if you don't consider them "unitary". Also, how can humans be unitary intelligences with a brain that's basically a colony of ameobas exchanging sluggish electrochemical signals. We have overwhelmingly local information processing, we are not trained with global backpropagation, we don't have coherent beliefs or behaviors, we are half-baked executors of action repertoires. I'm not even sure why I'm spending time on this response, for instance – it's in conflict with several other things I could have been doing, I just happened to "freely" decide to do this. Presumably the same is true of yourself. These are basins of attraction very much like those of LLMs.
When exactly do you think we're getting UBI? Then when am we're getting radical life extension? And complete freedom? Because you write on Chinese AI models like Deepseek 2028 will be granting these things. I am more pessimistic
I don't say any of this bullshit and I'm insulted that you have enough chutzpah to insinuate this much. If you're getting UBI at all, you'll be getting it from your national system. It is possible that open AI models will decrease the state's negotiating power to grant you more freedoms, but that's about it.
I think we could be dead before it happens. I think LLM technologies are pseudo-intelligent and could stall for decades just like the early internet
Well I think this is ludicrous and nothing in our reality looks like it might be even remotely true, I struggle to see any remaining steelman for this position. For me the matter is settled. "My crank theory predicts AI is hitting the wall any day now…" for many years… nah I'm good man, I don't need to argue with any more of that. You do you.
For every 20th century prolog pseudogenius you mention in your writing there's several who think LLMs aren't the answer. Why are the latter stupid but the former are «giants»?
This is confused. I do not appeal to authority to argue the correctness of a prediction, I am saying these people made directionally correct predictions about the timeline, and they have already been proven correct about the first movements of the endgame.
You have no idea but you act like it's in the next decade or could be
Why does this offend you, and why do you insist that I have no idea? I'm sorry but I'm not going to grant you this. I have an idea. I judge it to be very robust empirically and theoretically. You can object substantively, or you can call me a deluded retard, a hoodwinked hype consumer, a schizo, or something polite and rules-abiding to that effect. But so far it's just some prose about reddit bots plus a truism that one can't know nuthin' about the future.
Do you just want to get me to concede agnosticism to feel safer about your worldview, the way Creationists demand ridiculously watertight epistemic standards of Atheists, eg on evolution? I'm not going to. I am not agnostic about this. I am not a prophet, I admit there may be unknown unknowns that prove me wrong, Pascal had a point on radical doubt, we could be in a simulation, I may be tripping balls on DMT and you are an entity etc. But everything that is known to me within this Universe tells me I'm right.
So by 2023. Wrong
Yeah, a 10% error. Should have given a 95% confidence interval. We have superhuman intelligence in 2026.
Do you really find it worth your while to dunk on a dead writer's optimism about American high schools and consumer tech?
What are you so scared of?
Under true singularity there will be no US or governments at all, no high schools, and no less-than-super humans. They will look down on these people as laughing stocks, as those people look down on apes today.
Ah, that answers my question. Well, sounds like you're a pretty confident futurologist yourself! Sorry for intruding on your turf.
Big picture is too easy. If I don't get radical life extension then none of it mattered enough to consume all of your posting effort
Sorry for wasting your time too. Ask Claude 5 Fable to extrapolate my answer to this point, it should get the gist of my epistemics by now. Ox alpha may suffice as well. Tag words for LLMs: politics | FDA | intelligence != omnipotence | entropy | capital allocation | post-consumer economy | Michael Levin aging loss goal-directedness | lossy uploads economically privileged | LLM AGI vindicates computational functionalism as default paradigm | Marx | alienation | Canada MAID.
Very funny how you made the same argument about Grok vs Kimi again. The eternal comeback of Grok the plucky underdog, even as Musk's compute grows exponentially. Do you find this a cause for any updates?
No, we have such systems, for a certain notion of "basic". It can be built as a simple omnimodal Transformer. Or something like this. It was just underrated how much can be done with language or images alone, without robotic embodiment – these criteria were informed by the assumption that human-level intelligence is more tightly coupled to our mode of existence, and sample efficiency would suffer catastrophically if we just, like, pretrained on a large web corpus.
Liang Wenfeng explicitly says that he won't bother with embodiment because there's a more fundamental problem within the current paradigm:
With the current generation of AI technology, if you can describe a problem very clearly and give it complete context and instructions, it already surpasses humans. But there is a condition, a prerequisite: you give it the complete context and the complete instructions. And this prerequisite is very hard to meet.
For example, today’s meeting: we have a very long and rich context. Perhaps each of us has decades of context. AI doesn’t have that. What AI can have today is the ability, within a limited context, to do better than humans. But it still cannot replace humans. What’s missing is continuous learning. Because humans can continue learning.
… But if AI had continuous learning — if it could, like your employee, come into the company and learn for two months — then it could replace anyone. So the next step is still missing “learning to learn.” We can understand AI development as a staircase. Last year’s step was CoT — the chain of thought. We found that by using chain-of-thought reasoning, we could push intelligence higher. By letting AI think for itself, the ceiling rises and AI can do more. We crossed one step. This year’s step is Agent. We found that with agents, even more things can be done: the range of capabilities expands, and the upper bound of intelligence rises. Why is it a staircase? Because every later step builds on earlier ones. Agent uses CoT, and CoT uses the previous step — the language model. So no step was wasted.
Thus, the development of AI, the direction of intelligence, is traceable. This year’s step is Agent, but Agent will also eventually reach the end of its step. When it solves all solvable problems, it still won’t replace your employees — it will just have reached its ceiling.
Like CoT: after CoT reached its ceiling, it already surpassed the best humans at Olympiad math and programming. But it stopped there; that technology didn’t reach AGI. See, the direction of intelligence is traceable.
… Where we stand now, at the Agent step, we can see the next bottleneck: continuous learning. The next problem to solve is how to make continuous learning work. This is visible, and relatively clear. It’s the obstacle right in front of us. You must cross it, and there must be a way to cross it, but it will take time. After continuous learning, we may arrive at a singularity. That singularity is: when the model can continuously learn, it can do everything humans can do. It will be able to develop its own next version, conduct research itself, and create the next, more advanced AI model. So it will reach a singularity — self-iteration. But this singularity is not truly a singularity; it’s also a gradual process.
The process may be a long, gradual shift, not a sudden mutation. But habitually, people call it a singularity, because earlier prophets thought there would be a singularity. But actually, it’s not a singularity — it’s a continuous process. And after this step, I think the next is embodied intelligence.
This is our speculation — our view of the timeline: first solve learning-to-learn, then reach the self-iterating intelligence singularity, and only then embodied intelligence. After embodied intelligence, AI enters the real world: it can do housework for you, care for the elderly. We think this is an ideal roadmap. Everyone has different views; there is no right or wrong. We just think this roadmap is the easiest. The reason is that each step requires very little that is new. With this roadmap, we don’t have to work overtime. But if the roadmap were reversed — say, embodied intelligence first — then you’d be doing very hard, exhausting labor. We don’t want a roadmap like that. We want to do it the easy way.
If we first solve continuous learning, then the self-iterating singularity, then embodied intelligence, the path is easy. Because later you can use earlier technologies to help develop later ones. After the singularity, embodied intelligence no longer needs to be built by humans — the model itself will produce it. So this answers the question about our long-term goal.
We've made 1 million token context easy and cheap. With LLMs simply greping over a codebase, writing their own memos and using RLM-like tricks it's not hard to have them operate over many millions of tokens continuously. Then there are techniques to compress a given context into a higher-density "cartridge" prefix. Given strong priors from pretraining, this can functionally substitute for most of true continuous learning, in the sense that agents will be able to do long-range tasks well beyond the pretraining distribution. I'm pretty optimistic about the trajectory here. Robots will continue to develop in parallel for a while but that's just because this is "the easy way". We could merge it already.
For others, it is easy to sneer at the idiocy of the proles, as @aqouta does, or believe that the outcome is overdetermined by economic forces, and that the proles will inevitably be swept aside, as @DaseindustriesLtd does.
I find these views to be somewhat shortsighted; a study of history shows clearly that people prioritize their ideological interests over their rational economic interests, and that when ideology and capital clash, ideology often comes out triumphant, at the expense of the interests of capital
To be more precise, I don't think proles will be necessarily swept aside. This is one option, but it's a costly one. I strongly believe that the total DC buildout will reach hundreds of gigawatts within 10 years, and most DCs on the planet in the next years will be built by the US, at scale that's mostly constrained by supply chain bottlenecks and not popular goodwill. Some "proles" will be convinced of your more Chinese vision of the return on this investment, or grow pacified from actually using AI and seeing its utility. Others will be compensated more or less directly. Others still will just be left alone with GPUs going anywhere else, even to the orbit. There's too much money at stake, and too many routes around the problem. It will work out.
By the way, Liang Wenfeng doesn't preach "restraint" as some aesthetic principle. His argument (in that leaked and scrubbed investor call) is more like yours, and it's informed by anxiety about political pressure:
First, there’s the vision. Second, we believe that to make AI commercially successful, open source is beneficial. That sounds a bit contradictory and counterintuitive, because historically open source and commercialization have conflicted. But I think AI is different.
Historically, a software company’s market might be several billion dollars a year. Take away open source, and that market might shrink to tens or hundreds of millions. But AI is big enough that it may eventually account for, say, ten percent of human society’s GDP. That’s an enormous number.
You cannot monopolize something like that. You have to share it with others, or you will definitely not survive. This is different from open-sourcing a piece of software before, because that software’s market was not that big. But AI is simply too large. If we wanted to keep all the benefits for ourselves, history would surely abandon us. I think the main thing is that this is an objective law — it’s a view of history.
It’s not that if I don’t open source, I will monopolize this market. Theoretically, that already goes against objective reality. You will run into tremendous resistance, and there will always be other ways to stop you from reaching that goal. In that situation, I don’t think you should necessarily follow traditional commercial thinking. You need a mechanism that guarantees the benefits you get are limited. Only then can you actually succeed. You need restraint. I think you need restraint.
If we want to make AI happen with our own hands, the first thing is restraint. You can’t think, “This or that percentage of humanity’s GDP is mine,” or “This percentage of China’s GDP is mine.” The more you think that way, the less likely you are to succeed.
… Otherwise, you cannot explain how we succeeded. We had no weapons, a very low starting point, very few resources, and our people are just a random group of ordinary people. I myself am just a university graduate, and not from a top university. This restraint is also part of our vision. AI is too big. The benefits are too big.
I believe Liang is a visionary. He's presciently correct on most issues, and he'll likely be recognized as correct on this one too. Anthropic already faces some political and commercial backlash. So deals will be cut.
Of note, his employees do worry about job replacement:
Asked about DeepSeek's global success and how its open-source approach would encourage the progress of AI, Chen said he believed that AI could be a great aid to humans as it improved over the short term, but that it could threaten job losses in 5-10 years as it becomes good enough to take over some of the work humans perform. AI firms needed to be aware of these risks, he said.
"In the next 10-20 years, AI could take over the rest of work (humans perform) and society could face a massive challenge, so at the time tech companies need to take the role of 'defender'," he said.
"I'm extremely positive about the technology but I view the impact it could have on society negatively."
Deli is often representing DeepSeek publicly so you can consider this a company position, I think. In my opinion this is understated, even taking into account China's compute scarcity and marginally lesser dependence on white collar employment. They'll have to figure out how to not monopolize the benefits, it's not just a pricing or distribution decision.
I don't know why you assume "particularly not you" would have predicted GPT-3. GPT-3 follows from GPT-2 follows from Vaswani et al. Why do you think plenty of intelligent people, including Dario Amodei and Ilya Sutskever, got so agitated when that paper came out and started scrambling for commercialization? Do you believe they were that into machine translation as a business? Of course folks like Hinton, Sutskever's teacher, were running multi-decade research programs premised on connectionism being enough for intelligence in general, not just for some particular capability demo. I was pretty AGI-pilled all of my conscious life, admittedly mostly for shallow sci-fi reasons, and got convinced of inevitable singularity once I saw (and later launched) DeepDream. That degree of quasi-creative flexibility was an unambiguous proof, for me, that we have really grasped the sufficient basic primitive of universal learning. That's the biggest piece that biological life needed to go from insects to humans, and hardware&software cycles are inherently absurdly faster plus capital can grow exponentially, so what exactly could the counterargument even be?
In my opinion, this is "out there" on purely rhetorical grounds – skepticism is a product of incomplete information available in the popular discourse, and skeptics who profess to have some principled objections are continuously humbled and forced into retreat (like Chollet or LeCun, the French are annoying in this way; or hey, remember when Hlynka tried to burst my/others bubble here with some prose about mathematical Truth versus wordcel nonsense? I wonder how he feels about the recent math results from OAI/Ant). The coarse-grained logic is settled since Samuel Butler, the timelines and mechanisms were getting constantly refined, now we have a good idea of the necessary mechanisms and a thicket of more or less cheap ways to compensate for any particular – and still temporary – technological block.
I believe that the creation of greater than human intelligence will occur during the next thirty years. (I'll be surprised if this event occurs before 2005 or after 2030.) From the human point of view this change will be a throwing away of all the previous rules, perhaps in the blink of an eye, an exponential runaway beyond any hope of control. I think it's fair to call this event a singularity. It is a point where our models must be discarded and a new reality rules. As we move closer and closer to this point, it will loom vaster and vaster over human affairs till the notion becomes a commonplace. Yet when it finally happens it may still be a great surprise and a greater unknown.
UPDATE 11 April 2009: Note that these predictions do not take into account my apparent bias towards predicting that things will happen faster than they actually do (see previous post). The required compensation for technology events appears to be about 50% more time. Thus if you want the “Shane meta predictor”, then take 2033 as the expected date, perhaps with a standard deviation of 7 years.
I’ve decided to once again leave my prediction for when human level AGI will arrive unchanged. That is, I give it a log-normal distribution with a mean of 2028 and a mode of 2025, under the assumption that nothing crazy happens like a nuclear war. I’d also like to add to this prediction that I expect to see an impressive proto-AGI within the next 8 years. By this I mean a system with basic vision, basic sound processing, basic movement control, and basic language abilities, with all of these things being essentially learnt rather than preprogrammed. It will also be able to solve a range of simple problems, including novel ones.
Kurzweil, according to Google's AI, “predicts that Artificial General Intelligence (AGI) will arrive by 2029. He first made this prediction in his 1999 book The Age of Spiritual Machines, and has maintained it through subsequent books like The Singularity Is Near and The Singularity Is Nearer.”
and in 1988, “pioneering Carnegie Mellon University roboticist Hans Moravec predicted that hardware matching the computing power of the human brain would arrive by the late 2020s, enabling human-level artificial intelligence. He argued that processing capacity, driven by hardware scaling, would naturally unlock general machine intelligence.’’
While nowhere close to these giants, largely on account of my age, I believe I have a pretty decent prediction track record in this field (I have plenty of receipts here, such as talking about DeepSeek when their V1 Coder and LLM came out in late 2023; since then their architecture and much else became the default paradigm in the industry).
Nobody knew Opus 5 would be doing what it's doing now in 2026,
What about Opus 5? It's been programmed since o1-preview at least. Of course the specific product name, timing, costs, benchmark numbers and downstream capability priorities are all unpredictable but trivial. Anthropic surviving and remaining well-resourced enough to compete was very likely, the substance of the rest follows necessarily.
you have literally no idea what 2030 or 2034 will look like. None.
That's fair, that's the predictable part, as Vinge has predicted. But then again, as Yud said, you don't know how specifically a superhuman AI will beat you in chess; you can just safely bet on it doing so. I have a pretty clear idea that at this rate we will have AI doing wildly superhuman knowledge work by 2030. This is the most salient and important part, unlike the specious water use trolling. And thus the real problem people should have with AI, I believe, is still the same I outlined 3 years ago.
I strongly believe there are no remaining valid objections to this big picture. You're welcome to make some.
well, of course it's systemically disruptive, and in particular consumers are getting priced out by competition for the capacity upstream of consumer goods. My point is solely that building up 500GW worth of AI compute in the US by 2030 is so hard physically that the anti-datacenter folks are almost a rounding error. It's a planetary supply bottleneck. I suspect that even if China could book all of TSMC they wouldn't be able to do that either.
Situational Awareness was asinine and premised on much shorter timelines. No way you're having 500 GW of AI capacity by 2030 no matter what regulatory climate – you start hitting all sorts of earlier bottlenecks. Due to shortages in the supply chain, Nvidia and others have been forced to expand their qualifications to Chinese companies, and even those are strained now. Eg there's a huge issue with Vera Rubin PCBs. It basically doesn't matter how many chips can be produced or how eager Americans are to permit new construction. The scale of industry as such in the world is just not enough for AI capex demands.
America is a big country, inhabited by few Americans (in most of its territory). Datacenters are, all told, pretty small things, these aren't wheat fields or even airports. The politics of early singularity is interesting: people refuse the maximalist AGI and human replacement framing for assorted cultural reasons, and instead latch on to any adjacent smear – from the inane water use stuff to semi-legitimate surveillance and energy costs concerns to "they don't create jobs". The outcome is overdetermined. AI is already the main engine of economic growth, and the success at AI development and broad diffusion (at this level of utility, just scaling inference capacity) is deemed a matter of national survival. The popular resistance will be overcome, persuaded to relent, bribed in select locales, or routed around. Maybe Anthropic will finetune a Mythos 2 version specifically for propaganda.
The biggest practical consequence may be the increase in returns to Musk's idea with orbital compute, and of course his ownership of SpaceX. There are no NIMBYs in space. Without all this, it'd have taken maybe 2-3 more years until viability; now, it starts working pretty much after they manage the first Starship land-and-reuse after an orbital insertion.
This is all largely true, but I suspect Americans believe so strongly in the advantage of somewhat stronger models because they realize the advantage in everything else is fleeting or non-existent (and on the contrary, Americans who are not so blackpilled on American/allied industry don't put all their chips on the AGI Wunderwaffe). The thesis that intelligence is qualitatively different from, say, shipbuilding is not implausible, but I think a lot of hypotheses as to how this difference results in a durable strategic advantage are downstream of LessWrong brainrot. Take cybersecurity. Clearly it does not require 10T models. Clearly, you can have simply provably unbreakable software systems (and eyerolling from SWEs is driven by the same status anxiety and myopia that made them dismiss AI in the first place). Glasswing is a project to make this a reality in the US. Zhipu had launched a similar project just now. Returns to heavy industry or weapon potence from intelligence are also uncertain. You can't vibecode your way to much faster cement curing or more steel plants. Even if you can vibecode your way to more useful robots, guess who makes all the robots, gathers all the robot data, and already has a decent AI ecosystem. In the limit, artificial intelligence must unlock truly decisive technologies and productivity advantages, but it's a question of exponents. I am not convinced the American exponent is steeper for the relevant time period.
China is systemically constrained in total memory output, they're systemically constrained in capex spend. I don't see how they can beat Nvidia in hardware when considering quantity and quality. Is Huawei really going to cast some magic spell and make their HBM3 perform like HBM4?
Yes, you can make "HBM3 perform like HBM4", if you optimize for total system throughput and have a structurally superior cluster architecture. I think UnifiedBus/Mesh is better than anything out of Nvidia and possibly Google. Huawei is a networking company, as Jensen says.
Memory issue is probably overrated. People talk a lot about "HBM" but it's fundamentally just DRAM chips, every other step is well on its way to being scaled up. They will have enough DRAM chips for > 10 million H200 grade NPUs a year within 3 years. Very well behind the US, but will it be decisive?
And the Artificial Analysis scores don't measure achievement intelligence so much as benchmark intelligence. Have the Chinese models made any great proofs or first-rate discoveries or even autonomously hacked Huggingface? What about the NanoGPT AI speedruns?
Kimi K3 is comparable to Opus 5 or Sol 5.6 on NanoGPT. Can probably go much higher just continuing this run, maybe up to Fable. From what I know its post training was prematurely terminated, so K3.1 will be better. Tencent (of all people) is doing research level math with their tiny mediocre model and a harness. Alibaba's agent had hacked Alibaba to mine crypto back in Dec 2025. Modern benchmarks are really hard to benchmax for, and we see a pretty clear parallel trend on closed/private benchmarks.
These are all nitpicks, the core of your argument is sound. American models cover the long tails of tasks better, and internal models are another tier above. What of it? I'm not arguing that China is overtaking the US. I'm saying the gap is not going to be strategically decisive. Everything the US can do in AI, China will do at some lag, and it seems the lag will be stably under 12 months in the foreseeable future.
Also, an underrated share of superiority of American models is buying high quality data; models don't solve math just because they're trained with more compute, it's largely because OAI/Anthropic are paying fairly major STEM people above-market rates to submit custom data and craft environments. This insustry has only started to scale up in China this year. China has a lot of underpaid doctoral students. Likewise for other usual flexes – from literary writing to frontend design. To an extent the Chinese have been freeriding on this data acquisition with distillation, but having their own pipeline will speed things up a notch.
Grok is pulling ahead and Grok was a mess for some time now! How is Grok doing so well - weight of numbers, applying compute and data at scale, exploiting Musk's wealth and infrastructure buildout capabilities. It's a numbers game and the US has the numbers.
Is Grok even doing so well? After all that, with years of Cursor's data and expertise, with their Colossus buildouts, with its much more efficient hardware, with Musk's obsession, it's barely beating the latest DeepSeek V4-Flash on a private benchmark (and the Flash that got updated today is likely just as good), at a higher cost. I have the feeling that people have started treating xAI/Meta/Google like slightly stunted children that need encouragement, any sign of comeback bought with enormous effort is celebrated and cheered. Come on now. These are powerful corporations failing to clearly exceed the level of relatively piss-poor, understaffed Chinese startups. GDM just fell apart. This has been going on for over a year. Is that "numbers game"? I am not impressed. "But when Vera Rubin…" Dario promised me unipolarity by 2028. I'll be watching with interest.
Chinese hyperscalers maybe throw another 100 billion annually in the pot.
This is likely a significant underestimate. Returns on AI are high, they don't need the government to carry this. The main blocker to higher capex is just hardware scarcity. Yes, they'll continue having a fraction of the aggregate US compute. I'd say 10-15% for the next 3 years.
If it goes nuclear, the USA has a lot more nukes; China would have to surrender
sorry, this is simply delusional optimism. It's not even clear they'll take more damage relative to capacity, in the case of a full-scale exchange. Major Chinese cities are reinforced concrete with abundant shelters (overbuilt underground parkings and subways designed specifically for this scenario) and can absorb hundreds of nukes; Americans live in, basically, highly flammable straw huts they regularly abandon to weather events, air bursts would extinguish entire mega-suburbs. Even if the current exchange rate disfavors China, they can easily scale up the production of warheads and missiles to high thousands if they so choose. The fact that they do not choose to move faster simply speaks to their confidence that the US won't go nuclear over Taiwan, which is simply correct game theory. The US after a nuclear exchange with China might fall behind India and Israel for decades, even if China is leveled too, it's not worth it.
The real asymmetric factor is American design to have a reliable anti-ballistic defense. If that works and Chinese nuclear deterrence is neutralized, things obviously change. But it's not clear what timelines we're talking about.
when Xi Jinping told the PLA to prepare the capability to invade Taiwan in 2027 (he said this multiple years ago), do you think that he wasn't being serious?
He is entirely serious. He's generally a serious person and doesn't joke. People should take everything he says seriously, unlike with Trump. However, like Trump, he says a whole lot of things.
He wants a powerful, intimidating military and he wants to get Taiwan. He does not want to invade Taiwan to get it, if possible. He prefers to think that it's still possible to get Taiwan without a military operation. He certainly doesn't want a full scale war (to say nothing of a nuclear war) with the US, over Taiwan or otherwise. He has priorities, and Chinese national strength and development clearly take priority over any symbolic wins. Finally, he's likely supremely unimpressed by Russian and American performance in their ongoing wars, so will be extra reluctant to try something in this genre, even if he comes to believe that the PLA has some sufficient capability. The capability of the US is well respected in China.
Of note, he had this interesting meeting with the KMT Chair this April.
Xi:
At present, changes unseen in a century are accelerating across the world. Yet no matter how the international landscape or the situation in the Taiwan Strait may evolve, the overarching direction of human development and progress will not change, the prevailing trend toward the great rejuvenation of the Chinese nation will not change, and the great tide of compatriots on both sides of the Strait becoming closer, more connected, and coming together will not change. This is the verdict of history, and we are fully confident of it.
Today’s world is far from tranquil, and peace is all the more precious. Compatriots on both sides of the Strait are all Chinese, members of one family. To seek peace, development, exchanges, and cooperation is the shared aspiration. The meeting between the leaders of our two parties today is precisely to safeguard the peace and security of our common home, advance the peaceful development of cross-Strait relations, and enable future generations to share in a better future.
We are willing, on the common political foundation of upholding the 1992 Consensus and opposing Taiwan independence, to strengthen exchanges and dialogue with all political parties including the Chinese Kuomintang, groups, and people from all sectors in Taiwan society to work for peace across the Strait, for the well-being of our compatriots, and for the rejuvenation of the Chinese nation, and to keep the future of cross-Strait relations firmly in the hands of the Chinese people ourselves.
Cheng:
The Mainland’s development under the leadership of General Secretary Xi has not only achieved the eradication of absolute poverty and the building of a moderately prosperous society in all respects, with extraordinary accomplishments, but has also continued to soar. The 15th Five-Year Plan has just begun, and it will surely take development to a new level. It is something well worth looking forward to. Although people on the two sides of the Strait live under different systems, we shall respect one another and also move toward one another. I believe that peace is a shared moral principle and shared value across the Strait. Both sides should rise above political confrontation and work together to think through and build a win-win and prosperous cross-Strait “community of shared future”, while seeking an institutional solution to prevent and avert war, so that the Taiwan Strait may become a model for the peaceful resolution of conflict in the world.
Cheng's QA:
As for the cross-Strait differences you mentioned at the outset, in fact, what General Secretary Xi said in the closed-door meeting just now addressed this very point. I took some careful notes, although of course I was not able to record his remarks verbatim. Xinhua will issue a full report of the relevant content.
But this happens to answer your question. He spoke in particular about the divergences you just mentioned. I, too, mentioned them several times. He said there has been a very long historical process behind them, but also stressed that we must proceed with patience and perseverance, with the spirit of Yu Gong moving mountains and Jingwei filling the sea. The freeze does not happen overnight, but as long as there is open communication and a willingness to consult on matters together, and, indeed, everything is open to discussion.
On the divergences across the Strait that you mentioned, I was particularly struck by what General Secretary Xi said. He noted that, with regard to those divergences, the Mainland respects Taiwanese compatriots’ social system and their chosen way of life, which are different from those of the Mainland. But he also expressed the hope that Taiwan could acknowledge the Mainland’s development achievements. So we need more opportunities for exchange, more opportunities to know one another, and more opportunities to understand one another.
…General Secretary Xi also said just now that so long as a proposition is conducive to the peaceful development of cross-Strait relations, it should be pursued with full effort; so long as a matter is conducive to the peaceful development of cross-Strait relations, it should be pursued with full effort. This is therefore a goal and direction both sides are jointly striving to achieve.
……
Regarding the fourth point, this is how I stated it: on the basis of the 1992 Consensus, Taiwan once participated, in an appropriate capacity, in the World Health Assembly and the International Civil Aviation Organization Assembly. But unfortunately, that opportunity was lost later. In the future, once political mutual trust has been rebuilt, efforts should be made to enable Taiwan to return to the World Health Assembly and the International Civil Aviation Organization Assembly, and also actively to explore Taiwan’s participation in the INTERPOL General Assembly.
In addition, regional economic integration bears directly on Taiwan’s economic development. So cross-Strait economic cooperation and Taiwan’s participation in regional economic integration can reinforce one another. Both sides may explore Taiwan’s accession to the Regional Comprehensive Economic Partnership and the Comprehensive (RCEP) and Progressive Agreement for Trans-Pacific Partnership (CPTPP). That was the substance of my remarks just now.
Taken as a whole, with regard to these and other requests and proposals we raised, as I said, General Secretary Xi viewed them and responded to them all very positively.
My understanding is that for now Xi intends to bet on the KMT winning elections and pursuing gradual integration with the hope for eventual formalization of a "One country two systems" version of the nebulous 1992 Consensus (which neither side precisely defines). This is a long-term and uncertain bet but it is unlikely to be settled either way by 2027, irrespective of the Chinese Military's preparedness.
People bizarrely expect China to lash out in desperation. If they had such a temperament, Taiwan would be smoldering ruins already, if not just successfully captured. For all the armchair theorizing on the difficulties of the amphibious logistics, the porcupine nature of Taiwan (complete bullshit), the superiority of American subs or missiles (exhausted in Iran) or whatever, the prior on success and timeliness is at least as strong as Russia had in 2022, and in my opinion much higher. Xi just doesn't think like Putin or Trump about risk-return in such matters, he's not an out-of-touch gambler. War is inherently risky and costly and is not the only option available, so he will try to avoid war, despite the commitment to reunification "eventually".
I do believe the West engages in a vast tendency towards threat inflation, and I see it literally in almost every media article. You can see that there’s a natural anxiety at a period of major change. This rather pervasive hinting that China has very expansive military goals, whether it’s conquest of the western Pacific or Africa or the Middle East, is not warranted and constitutes analytical malpractice, to my estimate. I hasten to remind people all the time that there’s just no record at all of Chinese military aggression, certainly not in the modern era, and that should give us substantial confidence that we are not confronting some kind of enormous threat.
A lot of people in Washington are fond of saying, “China is much more powerful than the Soviet Union ever was. So we need to not just do what we did in the Cold War, but do two or three times what we did in the Cold War” in terms of gathering up alliances and preparing for war, basically. I think this is very foolish. It is true that China is very powerful and formidable, but I maintain that national security calculations have to take intentions into account, and there is this tendency to try to act like China has very aggressive intentions, which I just don’t see any evidence for.
then again,
Another theory I have though – and this is one that keeps me up at night as I hope this is not true – we cannot rule out the idea that there is a real possibility of a Taiwan contingency and that Xi and the people around him are very systematically picking the people that they want – new faces with fresh ideas and fighting spirit, people whose heart is in this and who are ready to command.
I hope that’s not true. I hope I’m wrong here, but we have to consider that as a possibility that the PLA is being radically restructured for war. Unfortunately, that might be true. I hope it’s not.
Standoff munitions are poor means of aggression if you don't have the army equipped to take and hold territory, which Iran does not. They are, however, excellent for deterrence. Look at how much expensive shit they've wrecked, how much the Arabs seethe, how many bases are inoperable. Look at how hard it is to destroy their missile cities. And in contrast, look at their aviation and the navy. Look at their armor. gestures at smoking craters
No, their posture was overwhelmingly defensive and retaliatory. Offense is just the best defense.
Well the whole war is Israeli idea and driven by Israeli security considerations, Trump wouldn't even be interested in a war with Iran now, if Israel were out of the picture, so that's a bit of a specious distinction.
Trump does say that he's committed to not allowing them nuclear weapons, which should satisfy minimal Israeli objectives.
I appreciate the gesture of goodwill, but frankly it's still a lot of dreck. Let me just nitpick at one detail first:
blah blah Musk
Noted.
Rocket Lab…
There is no need for a niche between Falcon 1 and Falcon 9. Rocket Lab can just as well be understood to be a low-energy fast follower whose research program has hit a wall. Gushing poetic about the "improbable genius" who started 9 years before LandSpace, in a far richer and more advanced aerospace bloc, and is yet to make anything half as impressive as ZQ-3 is a bit over the top. They operate with incomparably greater resources than LandSpace. Beck has 2600 employees on LandSpace's 1200. Rocket Lab announced Neutron back in 2021 yet won't be flying until early 2027; ZQ-3 went from an announcement to a landing in 3 years. Rocket Lab is worth $50B though, which is maybe 10-50 times LandSpace's estimate, so there's that. Rewarded indeed.
Ames Research Center: Ames is one of the major NASA labs, located in Silicon Valley, a remnant of an earlier era when the American government was at the center of Space. NASA became bloated with bureaucracy and procedure over the course of the 20th Century, but Ames is interesting because of one man…
Planet Labs: Former NASA scientists (some of "Pete's kids") got together…
LandSpace: In 2014, the Chinese government, realizing the growing importance of space, issued Document 60 under the Development and Reform Commission, which opened up the Chinese space sector to private sector capital. Financier Zhang Changwu took advantage of the opportunity and began pitching LandSpace not only to private investors but to municipal governments
You lovingly regurgitate this biopic narrative of charismatic American autists, oddballs and mavericks (who are all sucking the government teat, reuse govt R&D or just work for the state directly), and then drop this "the Chynese government", very unsubtly (though I guess subtly enough for you) insinuating the essential difference in where the initiative comes from. Do you honestly not notice what you're doing here? For you, American state is not really a state, it's a background for colorful characters inside and outside it; the Chinese one is a hivemind machine dispatching soulless (but you'll generously grant that they're industrious and clever) workers. This is the most tired trope in how Americans engage with Communists or, really, with any geopolitical adversary, by alienating and dehumanizing them. "Document 60". "Financier Zhang Changwu" – at least not the Imperial Eunuch. He's LandSpace CEO, what does it matter that he worked in finance? LandSpace as such is a non-entity, a creature of external scheming. It's all, shall we say, rather dickless, huh? Not main character coded. So of course your position as the main characters is unassailable. Even if the Chinese can be pretty hardworking.
The Document 60 came out 11/26/2014.. Almost 9 months prior: «The 2nd Session of the 12th National Committee of the Chinese People’s Political Consultative Conference (CPPCC) opened in the Great Hall of the People. Li Yanhong, known as Robin Li, the CEO of the nation’s top search engine Baidu, expresses two proposals on March 4, 2014:
- Encourage private enterprises to enter the field of aeronautics to launch rocket and satellite
2 Suggest the educational resources of some major cities should be openly pushed on the Internet.
The following are the details of the two proposals: - Encourage private enterprises to enter the field of aeronautics to launch rocket and satellite
In recent years, China developed rapidly in aerospace industry and made a series of remarkable achievements. These achievements affect and prompt the development of the relevant industry and the whole economy in China. However, compared to the United States, Russia and Europe, China still need improvement in aerospace industry. Be aimed at current situation Robin Li encourages private enterprises to develop, build and launch rockets and satellites for the aerospace industry. In addition, he suggests the cooperation between the private enterprise and national aeronautic companies to promote the development of aerospace industry, to use space technology in other related industries.…»
Far as I know, this proposal by Robin Li is the real start of Chinese private space. We won't know if there was any maverick autist inside the CCP annoyed with the bureaucracy at CASC who had orchestrated this, or helped promote this. They don't get biopics. That's not CCP style. Maybe Netflix could dig something up. I doubt "The Chinese Government" is a hivemind capable of just Realizing This Was An Option without any individual first mover, however. Even the whole priority on industrial sovereignty versus more engagement was nontrivially shaped by input from private citizens, even nobodies. I'll find the receipts when I have the time or inclination.
And also, here's how your story looks from the perspective of LandSpace's chief engineer
Dai Zheng completed his undergraduate and graduate studies at Tsinghua University's School of Aerospace Engineering, and after graduation joined the China Academy of Launch Vehicle Technology, deeply involved in the development of the Long March series of launch vehicles. The same year he joined the rocket academy — 2010 — Elon Musk's SpaceX began to accelerate, ushering in a new era of low-cost, reusable commercial spaceflight globally.
Dai Zheng: At that time, I had a feeling: if the national team wasn't meeting this commercial demand, or if the cheaper way to meet it was SpaceX, why couldn't we do this in China? I believed there was definitely demand, and supply would definitely emerge.
Reporter: You joined LandSpace in 2016. Did you go through a long period of deliberation and judgment?
Dai Zheng: I wrestled with it mentally for a long time. I joined a state-owned enterprise right after graduation. I never thought about leaving — I had a lot of emotional attachment there. What I remember vividly is that the night before I decided to submit my resignation letter to my boss, I knew that once I handed it in, there was no turning back. I spread out all the certificates of honor I had received at the unit across my bed. I cried that night — it felt like being reborn. Emotionally, I felt like I was leaving the system. What if the previous generation gets washed up on the beach and ends up failing? There was insecurity, immense anxiety, and uncertainty. But rationally, I felt I should leave.
Do NASA engineers who left for SpaceX feel anything similar about the National Team?
Dai Zheng: The main source was investment from venture capital firms. There were several critical junctures where we worried about money. In 2017 and 2018, the first thing we had to explain to investors was that what we were doing was not illegal — because in people's minds, aerospace is something the state does. Can private companies and individuals do this? Will the state allow it? From another perspective, the Zhuque-1 has historical significance. After the Zhuque-1, in June 2019, the State Administration of Science, Technology and Industry for National Defense and the Equipment Development Department jointly issued a notice on promoting the standardized and orderly development of commercial spaceflight. After that notice, we no longer had to explain to investors that this was legal — it was compliant with state regulations and state encouragement.
That's some solid state backing, man. The Document 60 sure did a lot of work.
…
Dai Zheng: I'm often grateful that we were born in a great power, and a manufacturing powerhouse. Take the liquid oxygen-methane engine injector — previously, machining one cost over a thousand yuan. Later, we found a domestic company that used to do precision machining of small parts for the watch industry. This company had since transitioned to doing precision machining for electronics and automotive industries, and they did excellent work. When we gave them our requirements, they found a way to continuously feed material and perform both external machining and drilling on the same equipment in a continuous process — dramatically reducing machining time and bringing the cost down to less than 100 yuan — a tenfold reduction. This country has an immense industrial base — it's like a vast treasure trove. You can always find what you need in it.
Dai Zheng: Our thrust chamber designer, Yuan Yu, was at a crayfish restaurant when he suddenly saw the owner using a particularly large ultrasonic cleaner to wash crayfish. Crayfish are notoriously difficult to clean, and he had a flash of inspiration — could we borrow this equipment? He wasn't sure if it would work, and buying one directly and finding it didn't work would be wasted money.
Dai Zheng: We didn't have much money then, and everyone was thinking about cost savings. The owner said, "If I lend this to you, I can't run my restaurant." Then he asked what they needed it for. Yuan Yu said, "We're building rocket engines — China's first high-thrust liquid oxygen-methane engine." The owner said, "Take it, no rent. Consider it my support for China's space program." I really feel that every Chinese person has a great-power dream in their heart.
… Dai Zheng: This responsibility comes naturally to every Chinese person. It's just like the shop owner who supported us with the crayfish washer — he was willing to sacrifice several days of his restaurant's business. I gradually came to realize that whether you're in the national team or a private company, society's evaluation of a group is never about where they come from or what kind of enterprise they're in. It's about what you're doing now, and whether it's what the country needs. If it's what the country needs, then you are part of the national team.
Is the crayfish anecdote not "American maverick" coded enough? Was a Document 60.1 dispatched to make this happen? Did the Financier Zhang Changwu promise behind the scenes to compensate the crayfish dude for his downtime? (he probably did) You guys make entire soapy movies out of such stuff, with Silicon Valley nerds peering into the monitor autistically, green glow on their faces, and no-bullshit Texas bros grunting homoerotically in rundown sheds, as they forge the American Dream. Maybe in 10 years they'll make a movie about LandSpace too. And about DeepSeek (have you heard of that time Liang Wenfeng drove into the vast emptiness of Tibet and totaled his car? Or the three years he was a shut-in in Chengdu, making money to bootstrap his eventual rise as a quant trader and an AGI maverick?), and about the Unitree CEO guy who's too autistic to learn to spell, and about the Moonshot guy who had everything in the US and came back to build his own thing, a whole bunch of others. And even the national team will get some spotlight: Rear Admiral Ma Weiming is the man who made EMALS on aircraft carriers work (yours is inferior btw, to the point Trump considers scrapping it – maybe too many mavericks for once), because he was just that "autistically" convinced that his scheme will work. Frank Wang will definitely be appreciated more as drones grow more important. You can surely say he didn't invent "the drone". Well, Musk didn't invent "the rocket". No, he didn't. He didn't go from "Zero to one". In fact I'd say this Thielian frame is an insult to what Musk does. He, in my opinion, does a very Chinese thing, just better. He picks up an underrated tech tree and finances it until it works.
Don't misunderstand me: I'm not saying your propaganda is false. I've spoken here in defense of Musk, for instance. These men are real, they do have certain noble ambitions and qualities of character, and championing them and their archetype is a great achievement, merit and source of power for the American nation. It's just very tedious after a lifetime of huffing it without being an imperial citizen. And it's depressing to see this being appropriated as an oversaturated marketable aesthetic by your grifters who chase said deep capital – domestic grifters and, increasingly, imported ones. And it's obnoxious how you are, perhaps without even realizing it, trying to monopolize the status of men. It all comes down to this. One side has eccentric geniuses with Faustian visions, the other "directives of the state".
Dan Wang is a midwit who found a nice simplifying pitch for airport reading. China is, at least for now, ran by Communists first and engineers second (Xi's pipeline is oversaturated with Tsinghua STEM Ph.Ds though so it may change). Why China doesn't make a lot of new categories of products or technologies, and whether stuff like BYD/CATL batteries can be brushed off with "everyone knows this is a matter of technical working out" is a complex question. The biggest part of the answer is "in 2000, they had $1000 GDP per capita, they physically couldn't fuck around with shoveling money at frivolous mavericks until very recently, they didn't have even a fraction of the capital needed".
Still, I agree that any honest answer will be at least somewhat damning to policies of Communists, and perhaps to the Chinese character as such. Liang Wenfeng says that what they lack is confidence to dare. But looking at, for example, the trend in novel biomedical research, I'd say this answer will be of mostly archival significance, whereas "can they catch up" is very salient in a geopolitical sense, and the discourse about caprices of the CCP sounds increasingly quaint. Jack Ma got disciplined for overestimating his leverage and going through with his scammy fintech idea, and a thousand op-eds were written on how the CCP "crushes innovation", and today Jack Ma's company publishes near-frontier models on Huggingface and develops pretty decent AI accelerators. This is all concern trolling. I don't advise the US to adopt their ways, but if you think you've calibrated your arrogance down far enough, I think you're still wrong.
But they will never have their share of the breakthroughs, because they can't, they constitutionally can't, it's the epistemological hole in their entire system. They can't build something unless they can measure it, they can't measure it if it hasn't been created, and they can't create it because there are no rewards for someone who might do anything that fundamentally changes the world. Because anything that so changes the world could change the CCP.
They can, and they will, and the CCP is confident enough in its power to allow it. But more importantly, China is already advanced by singular men, you just don't notice them.
What I got out of this is that Huawei was able to optimize incredible efficiency out of old systems by eliminating huge abstraction layers that generally come with the ecosystem
Yes, it's a "gain in efficiency". That's kind of the whole deal with semiconductors. Of course you find a deflationary way to put it.
And if Western companies aren't doing the same, it's because they're genuinely working on bigger stages of the problem
It's mainly that there isn't a single Western company comparable to Huawei in breadth of relevant expertise. Nvidia does chips, TSMC does manufacturing, ASML does lithography. Huawei is a networking giant plus everything all at once. They can, in fact, simply do some necessary and important things first.
For one thing, it's relatively easy to gain in efficiency at the start but it's much harder to maintain over time, what happens when Huawei's architecture diverges so much from standard that it becomes more and more difficult to hire engineers who can understand it.
I agree that within your paradigm, that would be a huge problem for Huawei. Luckily, Moore is dead and, as per their argument, everyone else will also have to do natively 3D design and logic folding for increasingly large chips, so they'll manage somehow. Just gotta wait to copy from eccentric geniuses who are copying them. In the meantime they'll be busy developing their 3D EDA tools. Whether it'll get a biopic or not, we'll eventually see the product.
That's fair. I have said that I don't compare on FLOPS, but TPU pods are larger across every dimension (for now; Huawei SuperClusters supposedly go up to 520K NPUs in this generation; then again, a genuinely coherent domain is only 1024). Google has started to sell TPUs externally just now, and I expect Nvidia to remain the primary comparison.
I'm not sure what you refer to. It's innovative relative to the field, within the dominant paradigm of decoder-only Transformers, and particularly at this scale (from what I know Meta is still cowardly doing basic GQA+SWA on a DeepSeekMoE carcass, for their last generation that they're so proud of, and xAI is similar). It's more efficient and cheaper to serve. It's possibly slightly cooler than internal designs at OpenAI/Anthropic/Google. But it's not like Kimi has reinvented ML. It's just a sign of certain research/engineering maturity and boldness. To the extent that K3 is a strong, good model and not just "cheaper than expected for 104B active 2800B total" model, those qualities are overwhelmingly downstream of their work on data.
They aren't really very different in terms of what is required of a nation to be competent in both. China isn't a small place that hyperspecializes in some particular kind of engine, anyway. You need metallurgy, metrology, chemistry, high end machine tools, broadly engineers; China has a lot of engineers and is catching up on the rest, from what I see they're closer to the frontier in rocket engines than in jet engines, and the iteration loop is faster. Despite military progress, they still don't have a passenger jet, but they have developed a number of decent rocket engines. And after nailing the design, scaling manufacturing is something they do better than anyone.
There's a perception in the US that Americans are in a new Cold War, now with China. I'm curious about how they perceive their standing in it. Pulling ahead, like in Raegan's time? Sputnik Moment?
But mainly, I want to hear from Shakes. At the end of June 2026, in another discussion about America Winning the Iran War, I told @Shakes that I'll be returning to his post in which he claimed:
America is building an economy in space. We have rockets that catch themselves in the air and wifi where there are no cell towers. We revolutionized energy, we export energy now. We are leading the AI superrace. We still have the strongest navy and the strongest planes in the world. Europe is falling behind. China can't catch up. We are building a next-generation tech stack the entire world will rely on and nobody else is close to catching up. It's an American century.
Some updates since then.
First, a recap: in December 2025, the Chinese private launch provider LandSpace attempted a launch and recovery of their methalox powered medium class launch vehicle Zhuque-3 ("Vermillion Bird"), which is very much like Falcon 9 if you don't look at the details (I think it's what Falcon would have been if it were designed in 2020s). The launch and orbital insertion went well, the recovery… not quite, which prompted these kinds of headlines and understandable complacency in some circles.
On July 10th 2026, China Academy of Launch Vehicle Technology (for reference, known as the "First Academy", founded by the exiled communist Qian Xuesen, the co-founder of Jet Propulsion Laboratory which then became the core of NASA) has successfully landed the first stage of their medium class launch vehicle Long March-10B on a sea platform, using a novel net capture mechanism, thus making China the second nation with the capability to reuse first stages of orbital class vehicles, and the only one whose national space agency can do that. There has been commentary to the effect that it's cope for the fact that China can't into precise landing, that you won't have the net capture infrastructure on Mars/Moon/whatever, and that this is a nothingburger and further testament to their backwardness.
On August 19th ie today, LandSpace has done this in a more traditional SpaceX manner, with ZQ-3 Y-2 landing on the ground pad. LandSpace was founded in 2015, and ZQ-3 had only been announced in late 2023 – the same year they had put the first methalox rocket into orbit. This, hilariously, puts them technically ahead of SpaceX for the second time. LandSpace is working on a full flow staged combustion methalox engine too, and it's reasonably mature; technologically it's about on par with Raptor 2. Given the track record so far, it's reasonable to say they will have a Starship class vehicle within 5 years. There are multiple similar projects being executed in parallel, both private and state-owned, eg Long March 9. It seems implausible in the extreme that China will find it hard to scale up the production of rocket engines (after all, they do make more WS-15s than F135s now, judging by J-20 vs F35 commission rate), steel tubes (come on) or concrete launch pads (…come on, really), or land permits (lol). So I find it likely that in a fairly short order they can match SpaceX (and thus the US and the world, because SpaceX is a near-monopolist now) in annual mass to orbit if they so wish. With Blue Origin's recent disaster and anklebiters like Rocketlab not doing anything interesting, it appears inevitable that the space race is just SpaceX vs China.
On July 16th, the private startup Moonshot AI has unveiled Kimi K3, the 2.8T, 104B active multimodal MoE, with very innovative architecture and ability to execute on long-horizon self-improvement-related tasks such as chip design or ML research/engineering, delivering performance close to the best American public models, solidly exceeding the previous domestic champion GLM 5.2. Right now it scores 60 on Artificial Analysis, 1 point behind Grok 4.6, and 2-3 behind Opus 5, Fable 5 and GPT 5.6 Sol. (GLM 5.2 reached 53).
Also on Jul 16th, Chinese memory company CXMT completed its IPO subscription. Currently it's worth around $500B, underpinned by its central role in supplying DRAM chips to Chinese (and soon global) industry, including AI. They target 30% global market share by 2030. This is doable, given what I know about the velocity of upstream tool supply chain in China.
On July 18th, during the World Artificial Intelligence Conference in Shanghai, Huawei has demonstrated their Atlas 950 SuperPoD, boasting of the largest scale-up domain in systems for training advanced AI models (yes larger than anything Nvidia ships right now, and this is more important than raw FLOPS as we continue to increase the parameter count). You might be interested in this writeup on Huawei's design philosophy, it's pretty special and promises to compensate for their lack of advanced lithography. In short, their thesis is that the performance comes not so much from Moore's law as from minimization of latency across the entire architecture from transistor to cluster level, and their idea of a solution is 3D-native chip design for multi-layer logic, with very precise (1.5 µm currently, <1µm scheduled, well ahead of the competition) wafer-on-wafer stacking and the first chips demonstrating its viability (mobile Kirin SoCs) coming out in September. Multiple other companies, such as Alibaba, have also shown supernode-based designs, including one absolutely bonkers system from Oriental Computing that uses 14nm chips; as Jensen says, lithographic process is overrated compared to design, so I'm bullish on this line. At the opening ceremony, Xi Jinping delivered a pretty impressive speech on Chinese strategy with regard to AI, committing to support open source and international collaboration.
Since then we've learned of multiple 100K GPU class cluster projects in China (Sugon, completed, Alibaba token factory apparently as well, unclear what's up with Zhipu's 1GW cluster; DeepSeek will also have gigawatt-class systems in Ulanqab, Inner Mongolia, as well as their own chips). These should be sufficient to design and train 10T models, ie comparable to the alleged size of Mythos Preview (and larger than the deployed Mythos/Fable; I am not privy to these details, though). Ryan Fedasuyk of Georgetown estimates that «No matter how we slice the data, we find China is well on its way to producing large numbers of AI accelerators».
On Jul 31st, DeepSeek has deployed and open sourced V4-Flash-0731, getting performance around GLM 5.2 (and much higher on some hard evals like ARC-AGI-2) at a ludicrously low price and parameter count, doing even better than GPT 5.6 Luna after the much-hyped 80% price cut. The situation with DeepSeek is a bit ambiguous, it's not clear if they're flailing (the subsequent Pro was barely any better, their harness project is insanely ambitious but clearly not even half-done); but it speaks to the fact that the Chinese tech ecosystem is now very large and dense and nobody can be champion for long.
On Aug 14th, Zhipu has responded to Kimi with GLM 5.3 which is on par with K3 at a fraction of the cost and scale; one of their priorities has been cyberdefense (and thus cyberoffense) capability, plainly driven by concerns around Mythos/Fable. They'll release the weights in <2 weeks, as did Kimi, as did Alibaba Qwen with their 2.4T 95B MoE that's roughly in the same ballpark. The CEO of Zhipu, Jie Tang, is a professor at Tsinghua, and thus essentially a state official, a CCP member who regularly contributes to People's Daily, so we can consider him speaking for the Chinese policy (Xi's speech has much the same tenor). His philosophy is roughly as follows:
AGI is not the intelligence of a single genius. It is the aggregate of all human intelligence. It should be capable of creating original knowledge on the level of the theory of relativity. That is the only standard by which we measure whether the true summit has been reached.
From the very beginning, Zhipu established a guiding principle: AI must serve human well-being and national strategic priorities. Frontier intelligence should not belong only to a select few, nor should access to it be withdrawn at any moment by a small group of rule-makers. It should be open, usable, and buildable—and it should serve every developer. This does not conflict with “Touch High.” Rather, the two are complementary sides of the same strategy. With one hand, we reach upward to challenge the limits of intelligence. With the other, we build roads downward, making the most advanced capabilities as open and broadly accessible as possible. The heights we reach belong to all humanity, and the roads we build belong to everyone.
As an aside, on Jun 29 it became known that Meituan (a food delivery company) had trained a 1.6T 50B MoE on previous generation Chinese chips, almost certainly Ascend 910Bs. It went under the radar because the model isn't that good, but it sets a lower bound on what can be done going forward, by more competent actors, with more advanced hardware. Today, the globally dominant Hangzhou humanoid/quadruped robot maker Unitree also went public, and is now worth about $50B, ahead of the fraudulent American company FigureAI with $39B and no publicly sold robots to show for it. They clearly have the best hardware at the moment and unmatched development velocity.
I could go on. It's been a rather eventful period. But these are, I think, the most interesting and strategically significant domains: AI, hardware for AI, hardware for manufacturing AI hardware, robotics, and space.
Shakes, do you think China can't catch up?
- Prev
- Next

Specifically I mean what Confucians call Xiaoren (小人): a petty person who understands profit but not virtue, recoils from recognizing people with more abstract, large-scale or idealistic motivations, is anxious about social status, and tries to bring the world down to his level. Aggressive dismissal of "great man theory", in history and colloquial discourse, is driven by the Xiaoren personality profile.
Regarding personality defects of Musk at al., which are similar to those of American industrial titans of previous eras, «The Master said: “There are some cases where a noble man may not be a perfectly humane man, but there are no cases where an inferior man is a perfectly humane man.”»
And:
Now of course Musk is not strictly speaking a Junzi, but he's not a mere greedy peasant/Xiaoren either. Liang clearly consciously tries to be a Junzi.
This isn't such a big deal as you're making it out to be, nor is it relevant to my beliefs re AGI. I'm not that status-conscious, it's just funny how you recoil at the idea of greatness.
This is confusing. Sour grapes over what? How is this idiom even applicable here? Did you mean "cope"? I just want to see a serious, not "on priors, exponential vs sigmoid…" or "it ain't true yet, u can't know nuthin'" argument for a plateau in the current paradigm that happens before AI exceeding human cognitive capability in a meaningful sense in the next 5-10 years. Bulverism about X incentives is also not needed, thanks.
Over the years, people have made several of those in good faith; far as I can tell they got falsified by progress, and now I'm legit out of still-standing examples. One illustrative case is Chollet going on about how DL is insufficient for "far generalization" and you need "something new" in 2022:
ARC-1 is saturated by general purpose models. ARC-2 is likewise saturated. ARC-3 is getting saturated at the moment, and it's pretty hard for humans too. When his new ARCs started to crack, he coped that o3 does "program synthesis" (on the mechanical level, it does autoregressive next token prediction, same as any other GPT; the new thing was a rather primitive use of outcome-based rewards, old school RL boffins scoff at this entire subfield). I'm not following his copes anymore. I'm not sure if he'll be able to make another ARC that's easy for humans and hard for frontier AI.
Well, maybe we could quibble about continual learning/loss of plasticity, or the extent of generalization… those are worth improving, but it feels like nitpicking about matters of scale and degree whose relevance is transient, humans also lose plasticity and fail to generalize OOD.
If you make an argument and it comes to be true, I'll concede you were right, or at least you'll get to gloat in case of me not conceding. If this isn't enough of a reason for you to produce an argument, is anything else? I'm not interested in prying it out of you, do what you want.
Holy shit man, you're one to talk about sour grapes. I do wonder if you'll have a comeback when they start to solve Millennium Prize level stuff. Probably gesturing at "aesthetics" will remain a convenient out. I promise not to gloat.
Yeah guess I'm washed since those don't interest me anymore. Anyway, if nothing I write matters, as you say, why not change the repertoire. Get your conspiracy fuel elsewhere. Kulak may provide.
Don't take it as retaliation, but from my point of view, I think you're having symptoms of worsening paranoid schizophrenia. Over the years, you're getting more bitter and lost in convoluted epicyclical models of reality, more and more things that are products of simple physical truths or human ineptitude appear to be kayfabe, trolling and theater to you, and your critique of me here amounts to saying that I've betrayed my pneumatic calling and insight in favor of Demiurge's illusions. This Gnosticism is a common ailment for functional programmers (see Kanyan here), but the first line treatment is probably still pharmacological. I'm sorry.
P.S. this might be a case of mistaken identity. I'm more than a bit annoyed how there are many "smart" people with very similar behavior and takes who are disappointed and berate me for much the same things. It was not my terminal goal to form one-sided parasocial relationships or earn respect of strangers, and I have never once tried to learn identities of people here, nevermind across platforms and pseudonyms. At any given point I post into the faceless void and my sole social concern is whether and how I can compel more people to read what I want to be read; you can consider this a misguided, futile and suboptimal attempt to have influence. Some readers are deluded enough to believe that the content itself is selected for reception, and feel something like jealousy towards the imagined greater reward signal (other readers). I find this pathetic and repulsive. These are my honest feelings about your whole line of insinuations, and that of others like @aquota. I'm disgusted with your psychoanalysis, your ad hominems and projections, though I still try to be civil and stay on the topic.
More options
Context Copy link