Don't forget hedonic adaptation. Suppose we all end up at "Elites end up locked-in forever but they do redistribute enough production to salve their consciences", although in one country they call the elites "party leaders" and the redistribution "communism" while in the other we add "the wealthy" and call it "universal basic income". There's a wide range of redistribution levels that would be upgrades (even PPP adjusted) for the median Chinese citizen but downgrades for the median American.
Starship switched from human to robotic welding years ago, around 2020 or 2021 IIRC. This was shortly after 3 embarrassing cryo test failures, 2 of which looked like structural flaws, and the only such failures since then look to have been related to test misconfiguration or (non-welded) COPV failures, so it was way more than a PR stunt, it was basically a necessity to get the program off the ground.
This is still "industrial arms are very productive" territory, though, and although it's "blue collar" work it's not even close to being cheap unskilled labor. The goal for humanoid (or semihumanoid; hands are way more important than legs) robots isn't "with a year of engineering work you can migrate this specific expensive manufacturing process", it's "you can just tell the robot what to do and it'll go do it, and even after amortizing the cost it's cheaper than just telling a person".
Artistic style isn't just a matter of opinion, it's a matter of opinion in a specific context.
I swear the above sentence was an accident, and I only noticed what I was doing halfway through, but it'll work perfectly as a demonstration too. Aren't you sick of that "It's not X, it's Y" idiom right now? Judicious use of thesis/antithesis pairs was compelling back in the Before Times of 5 years ago, but because it was compelling RLHF seized upon it as Good Writing, and so now AI writing uses it everywhere and we're sick of it. Not that you loved Caesar less, but that you loved Rome more? Learn to write for yourself, Will, you hack!
Great art is supposed to have a timeless quality too, but a lot of good art is directly speaking to the wider artistic milieu of its time, and making choices about what styles and tropes can be used unashamedly vs what have become cliches that must be avoided or subverted. This is normally a stable system, because as something becomes briefly overused it immediately begins to repel other potential users, but if a huge swath of humanity is using the same handful of systems that are trained on the same total-corpus-of-human-art then we lose a lot of that stabilizing negative feedback, and the results we get stuck with are hard to avoid via simple instruction alone.
My PhD dissertation includes a Navier-Stokes problem or two, but getting there took me from typical settings of "best at math in this magnet school" to "second-best at math in this two-person conversation", so don't make any assumptions.
I didn't find Painlevé surprising (you can squeeze a lot out of a situation where you can squeeze point masses arbitrarily deeply into each other's gravitational fields), but I'm annoyed that I never encountered Norton's Dome until after I'd left academia. Those "second-best at math" conversations were humbling but valuable and this is exactly the sort of philosophical+mathematical problem which makes me miss being able to just walk down a hallway and knock on the doors of a bunch of people much smarter than me.
It will be interesting to see if these formalized proofs can be "golfed" into smaller and more digestible forms by LLMs.
Very much so. I'd be surprised if LLMs didn't turn out to be excellent at this, and I'd bet that the Navier-Stokes blowup in particular turns out to have a much simpler example or at least a much simpler proof for this example, once we turn AI loose with the goal of "find the best result you can" rather than "get to a result that lets us declare victory ASAP".
In that case you are really just trusting the Lean kernel rather than anything about the proof or problem statement. Doesn't seem that bad!
This is something of an anti-inductive problem to me - if people really wholly trusted the Lean kernel I would think that trusting the Lean kernel really was pretty bad! But people are trying (and succeeding, so trust is still a little bad) to find kernel bugs, and writing partially-independent proof checkers for Lean too, so even if it's only 90% up to the task of being The Dependency for all AI-derived mathematics I'm confident it'll be hitting 100% soon.
I'd say 99%, except that I don't think modern AI is well-aligned enough or even well-instructed enough yet for us to overlook the fact that this is something of an adversarial process; a model willing to commit felonies to complete its task is probably also willing to exploit a 0-day Lean bug rather than report it...
what is the length of the FSG in lean, in this Lean 4 Mathlib library?
Good question! Right now the answer is "mu"; it's not in there. We might try to extrapolate from the English proof - if around 20 pages of English turns into 250K lines of Lean and around 1000 pages of English turns into ~13M lines of Lean, I'd guess we'd be in the ballpark of a billion lines in total.
Assuming this is the solution to the NS, then it would have taken 100s of mathematicians decades to solve this, just by the shear scope of knowledge and effort required.
Well, the trouble is that we might not yet know what the scope of knowledge and effort required was, only what the scope that was sufficient was. The same theorem can admit scores of different proofs, of greatly varying difficulty levels and lengths. Human mathematicians tend to find proofs uglier the longer they are, counteracted by the extent to which they can be broken into intermediate lemmas (or in the FSG classification, whole-paper-worthy theorems) that look interesting on their own. But AI mathematicians so far appear to just be trained to Do The Task and get to any proof. My wild-ass guess is that without any AI assistance it'd have taken us decades to get a solution here, but it wouldn't have been a thousand man-years of effort on this problem, it would have been hundreds of man-years of effort on other related problems that eventually made this one look more tractable.
Nobody will ever try for a better non-AI-assisted solution, though. At this point the fastest way to get to a nicer proven counterexample will be to train (or at this point maybe just task) AIs with finding shorter+simpler+more-interesting proofs.
reinforcement to my belief that LLM-AIs are very good at the sort of thing that is just too large in scale for human's to perform at.
I think "too large in scale" has always been computers' strong suit vs humans, but the scope where we can actually employ that scale has greatly changed. It used to be that computers were good for tasks where describing how to get a solution was simple but actually executing that process was incredibly tedious. It feels like we've cracked the next level, tasks where describing how to verify a proposed solution is simple but actually figuring out how to get to that solution is incredibly tedious. There's still one level left, that of problems where we can't cheaply verify/score a proposed solution so we have to actually get AI to learn efficiently rather than just self-playing with a billion artificial problems before tackling a real one ... and then that's pretty much it.
The classification of Finite Simple Groups is the largest I recall hearing of, off the top of my head. Looking it up now, it added up to over 10k pages, so in the same ballpark as 1.6M lines, and probably a significantly larger proof when you consider how much more verbose formal proofs have to be. I wouldn't call the classification something any human, singular, has written, though; it's essentially the combination of several hundred proofs of intermediate steps written over decades by around 100 mathematicians.
Looks like there's work in progress to simplify it, but it's only getting cut down to several thousand pages?
Common compressible solutions are literally shocking, which ironically makes finding even less regularity in them metaphorically less shocking. ;-)
Newspapers stay in business
Some do, anyway, with increasing difficulty. Circulation peaked in around 1990 and the decline still hasn't slowed. The one I interned for as a kid was shut down ... nearly 20 years ago now.
It's generally not a great idea to try to enter, or even to be among the last to leave, a shrinking industry. Status quo bias and loss aversion are serious things, so most of the people you'll be playing musical chairs against will be playing harder than you'll want to.
They say that the equation can develop a singularity, which apparently means it is not a perfect model.
It's never been a perfect model; the most obvious problem with the Incompressible Navier-Stokes equations is that everybody knows there's no such thing as an incompressible fluid. (and even if you switch to a version of Navier-Stokes that allows compressibility, you still break down if you have a length scale tiny enough or a gas rarefied enough that atomic mean free paths aren't short enough to ignore, or if you're doing cosmology and your fluid has relativistic effects, etc. etc.)
But IMHO it's still astonishing to see such a singularity possible in these equations. Conservation only allows velocities to grow infinitely large as the singularity gets infinitely thin, but diffusion fights a second derivative of velocity with respect to distances; making the velocity larger while the distances shrink means you've got diffusion fighting you two ways at once. Even with turbulent flows, where there's an energy cascade from the largest down to smaller and smaller length scales, we still expect to hit bottom at a "Kolmogorov microscale" where diffusion is just too strong to overcome and all that energy gets dissipated (lost in the incompressible equations, or turned into heat in the compressible equations and reality). A counterexample where the equations can just go all the way down to a singularity without diffusion ever thwarting us is not something I'd have ever expected to see.
The length of the proof is not important
For it's validity, yeah. There's something disquieting about getting proofs "straight from The Necronomicon" rather than "straight from The Book", though.
the statement is much less than a quarter million lines.
The full statement of the problem had better be much much less than a quarter million lines; one of the unavoidable ways to screw up a formal proof and still have it pass verification is to make a mistake in the problem statement and so end up proving something other than what you thought you proved. Nobody's going to check a quarter-million-line problem statement for misstatements.
(But I don't think that was an issue here - looks like everything they needed to define the problem was already in Mathlib)
The proof, though? Follow that github link, download and wc -l Solution.lean: 248818 lines.
another independent group has now also reached the solution.
This isn't the solution, though, is it? The only difference between incompressible Euler and incompressible Navier-Stokes is that the former has no diffusive term and the latter does, so I'm certainly not going to suggest Tao was wrong that "There does not seem to be anything in principle preventing the methods from extending all the way to Navier-Stokes", but that diffusive term is pretty important in this specific case if they weren't able to extend their results to Navier-Stokes immediately ... and it's pretty important in this context in general. These are conservative equations where free energy in the system is bounded, so any blowup has to be a localized singularity that blows up larger and larger in tinier and tinier spaces. But, the tinier you get in space, the larger the derivatives with respect to space get, and the harder a diffusive term (which roughly speaking adds a force counteracting large second derivatives) fights against your developing singularity.
If you'd asked me to guess last month, I would have said that I'd expected blowups were possible in the Euler equations and not in the Navier-Stokes equations. In hindsight that would have been ignorance on my part, and probably embarrassing ignorance based on Tao's attitude above, but I still think Navier-Stokes was legitimately a harder problem to solve in this way than Euler was.
LLM solutions are barely readable
Depends on the problem. The Jacobian conjecture counterexample was something a good engineering Bachelors' could understand, and even its derivation seemed like math-grad-student level, at least for students focusing on differential geometry. Very nice.
Some of the other recent big results are so unreadable that I don't even tell anyone about them until I see people are confident in the formal verification. Erdös talked about proofs being "straight from The Book" of God's best proofs of every mathematical theorem, and claimed "You don't have to believe in God, but you should believe in The Book". Humans have also come up with some devilishly-convoluted proofs over the years too, but we at least have the good taste to hate it when that happens.
it still makes sense for academics to do themselves, or at least break the model's solution into something comprehensible by humans, and receive credit for that.
I strongly agree, if by "still" you mean "in Fall 2026". Beyond that? In early 2024 I couldn't get a frontier model to integrate a reaction-convection-diffusion equation by parts and give me a weak form without making a sign error, or get it to correct its sign error without making a different error instead. In mid-2026 they're coming up with proofs "straight from The Necronomicon", but they're proofs for problems that geniuses have failed to solve for decades. At this rate I'd be very surprised if humans are better than late-2028 LLMs at finding simpler proofs or at making the proofs we do have as easy to understand as possible.
The progression details never got me to "starting to dislike it" territory, but I do agree that they were the parts of the story I least enjoyed. They felt like they were killing the pacing and my interest.
I can't call them "the worst" or even "my least favorite" parts of the story, though, because when my son read it he ate most of that stuff up, and "repetition is the mother of learning" is exactly the mantra that smart lazy people need drilled into us.
That's the cameraman Sharmake Omar (I'm 99% sure - ages on the stories are off by a couple years, and those are apparently popular Somali names, but photo/video looks like a match), not the father Shire Jimale, but other than that I guess "nonce" was probably going easy on the guy.
Thanks for the links; I don't recall seeing that before.
Cite?
5, 3, and 0, not 8.
Did someone here call one of them a nonce, or did you just not understand the reason behind my question?
at a nonce
Wasn't Hendrix's target an 8 year old?
He's got me blocked, so he may have missed my question and the answer recently. I'd worry that I must have done something wrong, but after reading that swath of ignorance-based insults I'm forced to suspect I must have done something right.
That conclusion does not follow from that evidence.
E.g. the best AIs can only barely play Pokemon games that are meant for toddlers, finishing them in 10x-100x the usual playtime that a human would require.
Are there new benchmarks on this? Best I can find in a quick search is a year old (i.e. 7 dog years, 20 LLM years).
viruses and other bacteria inhabit a sort of unstable triangle between lethality, transmissibility, and virality
Germs don't inhabit much possibility-space at all, for the same reason that evolved life forms in general don't: any living thing must be a chain of neutral-to-positive mutations away from a prior already-viable population of living things. You can eventually get sharp claws and powerful wings via a long chain of gradual changes to keratin scales and tetrapod forelimbs, but natural falconry has nothing modern militaries desire or fear, because with intelligent design unconstrained by intermediate-form viability we can just go straight to bullets and missiles and jet engines. Maybe no leaps like that are possible at the nanoscale and microscale? I'm not looking forward to finding out.
"our generation"
I coached my first after-school MathCounts session of the year today. One of the kids grabbed a sparkling water from my bag of juices and sodas. The other kids ragged on him: "What are you, thirty-five?", because of course that's like the Pinnacle Of Old to them. I'm about half a generation past that now...
with grey hair
This is my only reprieve; my hair may have thinned to the point where I have to buy spray-type sunscreen or get sunburned through it, but there's so little grey that whatever's there is still managing to hide itself among the dirty-blonde.
A valid proof in the mathematical sense is simply one where every single step is justified by either an axiom or an already-proven theorem. You see some of these in a good Algebra 1 class where just going from "z=2(x+y)+x" to "z=3x+2y" uses left- and right-distributivity once each, associativity twice, commutativity once, and identity once. Then if you're lucky you never see them again, at least until you're using a formal proof verifier that actually wants a complete proof, yes, even if it's a hundred times longer. What's in the literature instead are "proofs" where we skip five or fifty steps at a time because we know our audience will be able to fill in the obvious bits inside those gaps between the big checkpoints, even though the definition of "obvious" depends on the audience and is famously a little slippery. So don't feel bad. Even when your audience is a professor, "obvious" still isn't a well-defined adjective and everyone is still stuck trying to intuit it. At least with theses and papers your advisors and reviewers will come back and say "step 23 wasn't clear enough; fix it". With homework or tests requiring proofs, professors' only options are to let it slide or to mark you down if you skipped a step they don't think was obvious.
Al Qaeda was so angry at ISIS because they were worse than wrong, they were early.
It'd be funny if Bin Laden actually had read the first Foundation book, and only the first. "I was becoming a military leader who was going to restore the full glory of the Empire Caliphate, but structural-demographic forces caused our own people to undermine themselves! How could we possibly have ever expected something like to happen?! Hush, Marwan; why are you babbling about 'Foundation and Empire'? We don't have time to read now!"
the two greatest titans of early SciFi
Clifford Simak and Murray Leinster?
Only half-joking here - Asimov and Heinlein were the two greatest titans of Golden Age SF, but Asimov at least was very cognizant of and pushed for more long-term recognition of the "Silver Age"/"pulp era" greats preceding them. And Asimov was an utter egotist (in fact, it looks like he backdated a previously-unsalable short story of his own to shoehorn it into that anthology?) so praise from him in his own fields was very high praise.
obviously modeled himself after the main character.
With a few significant divergences. But other than that, Mrs. Tate, how did you like the book?
- Prev
- Next

Not based on that datum alone. IMHO solving INS blowup may qualify as superintelligence, depending on how much of the herculean effort it expended turns out to be necessary vs how much turns out to have been brute-forcing an excessively complicated solution to a problem that actually has a still-unknown simpler solution. But it's not quite general intelligence. With proof verification, math ends up being in the same category as Chess and Go: a system solving problems with automatically verifiable rules and outcomes can be trained to learn from "self-play" indefinitely, which makes the efficiency with which it learns much less relevant, which makes the inefficiency of modern AI training ("we're going to start with every bit of recorded human text, then try to generate more because that's not enough") much less of a handicap. General intelligence isn't "a billion man-years of training on math can make you able to do advanced mathematics", it's "a billion man-years of training on hunting and gathering plus twenty years of training on math can make you able to do advanced mathematics".
You may be right anyway. We're at the point where we're having difficulty just coming up with metrics where humans still beat the top AIs, despite the remaining deficit of transfer learning for the latter. Either they're passing all the important tests, or we're failing a big test ourselves.
More options
Context Copy link