Rendered at 22:07:29 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
no-name-here 1 days ago [-]
Real title includes:
> Unsolved Problem by Fields Medalist Breached by Two High School Students with AI
My title recommendation: Fields Medalist Problem Solved With AI
They used AI for “computation, proof idea generation, and editing assistance”.
It’s a bit odd how they list the AIs used - “Claude Opus 5, Anthropic and ChatGPT Sol5.6 were used for calculations, proof ideas, and
editorial assistance.”
random3 1 days ago [-]
note HEAVILY
> The students heavily utilized AI assistants, specifically Claude Opus 5 and GPT-5.6 Sol, for computational exploration, proof idea generation, and editing.
this is a bit like Enhanced Olympics(https://www.enhanced.com), except that you have someone else compete for you.
My issue is that it makes it hard to distinguish real insight/work etc. from effectively null one.
An old instance of the same issue was with what was called "script kiddie" back in the 90-00s
gus_massa 1 days ago [-]
For short problems the LLM are just too good now, but for long problems they still get in trouble.
It's like bicycle or F1 race. It goes faster, but you still have to steer the boat to reach somewhere. (Or probably something in between, like a motorcycle race.) Also, the "kids" were guide by a postdoc, not completely on their own.
karmakaze 1 days ago [-]
Difference being anyone can be a "script kiddie". I don't think anyone could direct an AI to proofs like this one.
dist-epoch 24 hours ago [-]
> Jarred Sumner, an Anthropic staff member (and non-mathematician), prompted Claude to "take a real stab" at the hypothesis itself, leaving the mathematical choices from there up to the model. ... Throughout this process, Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of "keep going" or "believe in yourself"). This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
It means that the solution is merely an interpolation of existing work and not fundamentally innovative as AI can only regurgitate, never creating something new.
What I'm curious about is how far they would've gotten without the postdoc.
srean 1 days ago [-]
This is a common pattern ahead of admissions season. Now admissions are mostly done I suppose.
htrp 1 days ago [-]
the fact that you have to encourage these models and tell them that they can solve these problems and warm up on easier problems seems to indicate that there's something to AI pairing above and beyond prompt
asolove 1 days ago [-]
They are relying on training data of humans talking about how hard these problems are. Same way an un-reminded Claude gives estimates for work that are as if a human is doing it by hand, but then will drop them by 20x if you remind them it’s going to do the work.
bananaflag 1 days ago [-]
Yeah it s because they have a strong prior on unsolved problems being unsolvable.
Once the idea of AI routinely solving conjectures enters the training data this encouragement will disappear like 2023-era prompt engineering did.
arscan 1 days ago [-]
I would think this type of behavior (ugh, or dare I say default mindset) by consumer-facing LLMs will always be desirable for ‘hard’ problems (things previously unsolved) because it’s a bit like having saftey mechanisms in place to prevent hallucinations for users incapable of verifying correctness of the output. You’ve got to do a little work to prove you understand that it’s hard but it’s still something the LLM might be able to accomplish.
rowanG077 1 days ago [-]
I think this is also a healthy mindset for a person to have in many cases. If my boss asks me "go solve P = NP", as extreme example, I would also give some pushback.
skybrian 1 days ago [-]
It will likely burn a lot of tokens, take a long time, and might not work, so hopefully they'll still ask if you really want to spend the money on the attempt.
empath75 1 days ago [-]
I spent about a month walking through proving something with Claude a few months ago and it _constantly_ told me that it was impossible and I should stop working on it, right up until it proved it.
dist-epoch 24 hours ago [-]
One way they improved the hallucination problem was basically training the models to refuse to do or say something if they are not very sure they can do it. As a side effect, they refuse to work on problems they know are extremely hard.
Aargau 1 days ago [-]
Working backwards from the Navier-Stokes solution, I was able to walk Astra through the path used to solve it. It took some formulation, starting with the problem, then challenging it to look closer at the specific path, and iterating when it got stuck, it was able to reach the solution.
Xcelerate 1 days ago [-]
I sort of wonder if effectively using AI to solve math problems is a skill in its own right, distinct from traditional mathematical skills. I don’t just mean “prompt engineering” either. More so figuring out how to combine agents with other tools and approaches in an effective way.
firesteelrain 1 days ago [-]
Agree this is a broad statement for any use of AI for any purpose. Open ended prompts / requests are bound to lead to misleading results, hallucinated responses, waste of tokens, etc
breezybottom 1 days ago [-]
You still have to be able to verify the solution to say that it's solved, so I'd say no.
perching_aix 1 days ago [-]
I'm not sure that makes sense? Having the expertise to verify the completion of a given task is a usually necessary but not sufficient requirement to what they're describing. I don't think they even disagree.
You seem to be imagining completely independent areas of competence, but I don't think that's a reasonable interpretation of what they wrote.
stabbles 1 days ago [-]
If it is a skill, it's something that can be learned by both humans and LLMs.
threethirtytwo 1 days ago [-]
I guess the better question is, can this skill be learned without being an extreme expert in mathematics? Can I be a sort of intelligence highschool student (not incredibly exceptional at all) and have the LLM teach me the required math for a specific problem and guide it towards a solution?
I think the answer to this question is becoming, in general, a big yes.
LLMs can arrive at a solution via two paths. The first one is via heavy guidance by an adept expert in the domain. The other path is via brute force... multiple agents (the more the faster it can arrive at a solution). The later path is what enables anybody to do this.
empath75 1 days ago [-]
IMO, it _currently_ requires a lot of skill because it will frequently take wrong turns and dead ends and needs suggestions and steering to get there. You do need to understand what it is doing at least a little bit and to understand the general landscape of the problem, and to at least have a sense of whether and why the problem is tractable at all.
catpower 1 days ago [-]
Great short term achievement but humanity is better served by these kids doing it without AI. The ideas have to come from the next generation (eg would we be worried if a 5 year old wrote the great American novel with AI or not?)
skybrian 1 days ago [-]
I think it's more like how anyone learning to play chess is going to use chess engines, but you can't learn to play by always asking the chess engine for the answer.
I expect that all mathematicians are going to be working with power tools, so they might as well learn about that. They will still need to do math exercises by hand to learn the material.
measurablefunc 1 days ago [-]
AI is not going away so people will have to get used to it just like they will have to get used to it in software engineering. The alternative is being less productive than people who are happy to use AI to write software & do mathematical research.
adamddev1 1 days ago [-]
A similar line of thought is used to justify corruption:
"Corruption is not going away so people will have to get used to it. The alternative is making less money than people who are happy to take bribes."
btilly 1 days ago [-]
The fact that a similar argument can be made does not mean that it is wrong.
Socrates famously made the opposing side of the argument against writing. Which is why we mostly know of him through Plato, who did believe in writing.
And to your corruption example. If you live in a society where corruption is normal and expected, you will be worse off if you are unwilling to be corrupt. It is indeed a local optima. But we are all, of course, better off if we live in a society where corruption is punished. To me, the worst thing about modern US politics, is that it's encouraging us to see ourselves as living in a world where corruption exists and is tolerated.
selectodude 1 days ago [-]
I can’t run a corruption racket in my basement.
measurablefunc 1 days ago [-]
People used to complain about compilers as well.
adamddev1 1 days ago [-]
Compilers are deterministic.
measurablefunc 1 days ago [-]
So are AI.
adamddev1 1 days ago [-]
LLMs are probabilistic, not deterministic.
measurablefunc 1 days ago [-]
That is a very common misconception.
oasisaimlessly 1 days ago [-]
It is still an unknown whether an unqualified increase in productivity in the long-term for software engineering is a given.
qarl 1 days ago [-]
Really? OK - then let me stand up and vouch that I am at least 20X more productive with AI.
adamddev1 1 days ago [-]
Would I be wrong to assume that you are building end-user applications?
If people use AI for libraries, OSs, and mission critical software, the apparent productivity gains would have to be weighed against the reliability and performance hits that bubble up to the things that are built on them and rely on them.
qarl 1 days ago [-]
In my experience - a robust testing harness will get you the safety you need. And most software you describe has such testing.
I think the Bun port is a great example where testing enabled a very successful implementation. (Both the original tests themselves and runtime comparisons to the previous implementation.)
adamddev1 1 days ago [-]
Testing is an extremely inadequate measure of reliability and robustness.
qarl 1 days ago [-]
You understand that robust testing can, should, and often does test directly for reliability and robustness, right?
adamddev1 1 days ago [-]
Yes of course, but you also understand that as Dijkstra said "tests cannot show the absence of bugs."
Tests can show you problems, if you can find them, but they cannot show that there are no problems. Property based testing or fuzzing gets your more coverage, and is a good step, but it is still nothing compared to proving things or understanding how something is built and that it is solid. Testing works towards checking for reliability and robustness, but often it's only 10% (?) of the job.
qarl 1 days ago [-]
Well, yes, I understand that testing does not create a provably correct solution. But I'd love to hear the source of your "extremely inadequate" or your "10%" claims. I mean - there is a reason why it's used extensively in software engineering - right? Or don't you see value in that, either?
I'm curious - is there any data on the Bun port error rate? I think that would be very indicative of how successful or not the testing is.
tom_ 1 days ago [-]
In the long term, it'll be at least be the year 2030. Let's not get ahead of ourselves.
dermacentor 1 days ago [-]
In the long term.
qarl 1 days ago [-]
Fair enough - it will take some time before the science comes back.
But I feel the need to point out - the goalposts for "does AI work" shift daily.
vouaobrasil 1 days ago [-]
But isn't it a horrible thing that what you just described (being forced by the prisoner's dilemma), is what defines progress these days?
measurablefunc 1 days ago [-]
There is nothing I can do about that. Investors & shareholders believe that AI is the future so that's where all the money is going.
vouaobrasil 1 days ago [-]
I know. I just thought it was kind of a sad state of affairs....
measurablefunc 18 hours ago [-]
I agree but that's the reality.
thrill 1 days ago [-]
Haven't you heard? Fields Medalists believe it only counts when they do the solving - not the unwashed.
1299348 1 days ago [-]
Not peer reviewed. UCLA seems to be full of AI boosters who perform circus tricks.
Founderarcstone 1 days ago [-]
Great to see high school students getting attention on this.
robotpepi 1 days ago [-]
i wonder if anyone is going to read that.
esafak 1 days ago [-]
This looks like a graduate-level proof. Did these high school students really understand what they were doing? I'd like to see what they have to say.
> Unsolved Problem by Fields Medalist Breached by Two High School Students with AI
My title recommendation: Fields Medalist Problem Solved With AI
They used AI for “computation, proof idea generation, and editing assistance”.
It’s a bit odd how they list the AIs used - “Claude Opus 5, Anthropic and ChatGPT Sol5.6 were used for calculations, proof ideas, and editorial assistance.”
> The students heavily utilized AI assistants, specifically Claude Opus 5 and GPT-5.6 Sol, for computational exploration, proof idea generation, and editing.
this is a bit like Enhanced Olympics(https://www.enhanced.com), except that you have someone else compete for you.
My issue is that it makes it hard to distinguish real insight/work etc. from effectively null one.
An old instance of the same issue was with what was called "script kiddie" back in the 90-00s
It's like bicycle or F1 race. It goes faster, but you still have to steer the boat to reach somewhere. (Or probably something in between, like a motorcycle race.) Also, the "kids" were guide by a postdoc, not completely on their own.
https://www.anthropic.com/research/riemann-zeta
Title: BOUNDED RATIOS FOR LORENTZIAN POLYNOMIALS
https://arxiv.org/pdf/2609.05341
Once the idea of AI routinely solving conjectures enters the training data this encouragement will disappear like 2023-era prompt engineering did.
You seem to be imagining completely independent areas of competence, but I don't think that's a reasonable interpretation of what they wrote.
I think the answer to this question is becoming, in general, a big yes.
LLMs can arrive at a solution via two paths. The first one is via heavy guidance by an adept expert in the domain. The other path is via brute force... multiple agents (the more the faster it can arrive at a solution). The later path is what enables anybody to do this.
I expect that all mathematicians are going to be working with power tools, so they might as well learn about that. They will still need to do math exercises by hand to learn the material.
"Corruption is not going away so people will have to get used to it. The alternative is making less money than people who are happy to take bribes."
Socrates famously made the opposing side of the argument against writing. Which is why we mostly know of him through Plato, who did believe in writing.
And to your corruption example. If you live in a society where corruption is normal and expected, you will be worse off if you are unwilling to be corrupt. It is indeed a local optima. But we are all, of course, better off if we live in a society where corruption is punished. To me, the worst thing about modern US politics, is that it's encouraging us to see ourselves as living in a world where corruption exists and is tolerated.
If people use AI for libraries, OSs, and mission critical software, the apparent productivity gains would have to be weighed against the reliability and performance hits that bubble up to the things that are built on them and rely on them.
I think the Bun port is a great example where testing enabled a very successful implementation. (Both the original tests themselves and runtime comparisons to the previous implementation.)
Tests can show you problems, if you can find them, but they cannot show that there are no problems. Property based testing or fuzzing gets your more coverage, and is a good step, but it is still nothing compared to proving things or understanding how something is built and that it is solid. Testing works towards checking for reliability and robustness, but often it's only 10% (?) of the job.
I'm curious - is there any data on the Bun port error rate? I think that would be very indicative of how successful or not the testing is.
But I feel the need to point out - the goalposts for "does AI work" shift daily.
https://arxiv.org/pdf/2609.05341