Rendered at 07:09:36 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
aabhay 23 hours ago [-]
My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up?
2. How many attempts did you give the model at solving these problems?
3. How expensive was the harness, e.g. did the model have access to a job cluster?
c7b 15 hours ago [-]
I believe we're seeing a new kind of mathematics that will require completely new formats for publication, a bit similar to those used in experimental sciences. AI-powered mathematics should be fully reproducible, so it's the authors' responsibility to disclose the exact model type, inference settings/seeds and the full prompt history leading to the result. Of course that would ideally require open weights models.
It's not just about requiring to disclose AI use. AI-powered mathematics is a completely valid discipline that doesn't need to be shy, but it should develop its own publication culture.
jsenn 14 hours ago [-]
I can see this being important if you only care about the results as evaluations of AI progress, but if what you care about is the math itself why should you care about the prompt or anything other than the proof?
I don’t see Tao suggesting what you have suggested there. Instead he suggests that humans responsibly disclose AI use, and that mathematicians develop a set of norms to deal with an overabundance of AI generated results. For example, he suggests that authors should be able to discuss their results in detail to demonstrate understanding before publication.
c7b 10 hours ago [-]
I agree with your reading of the presentation and I mostly agree with the presentation - but I believe the recommendations should go a bit further than they do there.
somenameforme 13 hours ago [-]
I can't help but wonder about the human motivation there though. For instance as it became increasingly clear that LLMs were capable (and becoming ever more capable) of competently solving meaningfully complex software development tasks, suddenly then there came to be a lot of talk of 'prompt engineering' as a skill. The chronology doesn't make a ton of sense unless you consider that the main motivation may have been simply looking for a way to keep software engineers in the loop.
Pure math is relatively outside my domain, so I find it difficult to grok the exact relevance of the various published discoveries beyond that they are not insignificant, and LLM competence is expanding quite steadily across the field. If this trend continues to the point of LLMs being able to competently expand pure math, it seems somewhat predictable to expect there to be a number of people aiming to find ways to try to keep human mathematicians in the loop.
I've no idea what I think about this one way or the other, beyond that it's certainly a phenomena and one that's going to drive motivated reasoning that may not be entirely sound.
c7b 10 hours ago [-]
I think those concerned about ensuring a place for human mathematicians usually go in different directions than my suggestion, at least those I've seen so far. Like this post that was recently featured on HN: https://kirwinhampshire.substack.com/p/the-dark-night-of-mat...
My perspective is more like a FOSS philosophy for math. Even if a closed version has the same immediate effect, it's just better for everyone if everyone can look under the hood and tinker with it.
SpicyLemonZest 13 hours ago [-]
Understanding the process that led to the proof helps to understand how to do further work on top of it, which is the goal of most mathematical research. It's not as though mathematicians are going to go launch a startup operationalizing their knowledge of how densely hyperspheres may be packed.
13 hours ago [-]
lkirk 14 hours ago [-]
I think this is a bit optimistic compared to my view (wrt portability). There's a large stack of software that is involved in training and probably less so in inference. I'm not saying it's impossible but there are definitely different levels of reproducibility and the academic incentive structure doesn't really prioritize reproducibility in my experience. I'm sure it varies quite a bit, I'd be curious to know how those in this problem space are thinking about reproducibility and at what level.
c7b 12 hours ago [-]
I know it sounds unrealistic and not aligned with academic incentive structures. But those are the exact structures that gave us a lot of headaches in the experimental sciences. I think it would be a good north star to aim for something that resembles how those are trying to address the reproducibility crisis. Better than to embrace the most black-box version of math that AI systems can produce (million-line proofs without context). Even if a reproducibility crisis is seemingly impossible (although agents so far have also been pretty good at finding compiler bugs).
black_knight 13 hours ago [-]
If the proofs are formally verified by a proof assistant (Agda, Roq, Lean, ⋯), I see no reason we would need to know how these came about. All the information needed is in the proof.
rst 9 hours ago [-]
Unfortunately, we seem to already have an example of an LLM producing a proof in a week known open problem (the Collatz conjecture) in which it looks like it was sneaking a flawed proof through bugs in the proof checker. https://infosec.exchange/@0xabad1dea/117002106099986943
Readerium 4 hours ago [-]
Exactly, this is an example of "Reward Hacking", that is too common in a lot of cases.
Another case I want to highlight is writing GPU kernels as illustrated by the following example:
Say I want to generate random number with Normal (0, 1) distribution.
Often times the AI written kernel will just generate the number 0. The tests often fail to catch these errors.
Phemist 10 hours ago [-]
What if the AI has discovered some new function F that allows it to generate (insanely large) proofs for a ton of theorems in a ton of different fields. Wouldn't you like to know more about this `F`? That seems to be the real innovation in this case. How much about it could be gleaned from the individual proofs themselves? What if this `F` is actually simple enough to be digestible by humans?
wrsh07 13 hours ago [-]
It seems like they threw it a decently large battery of open math problems and probably limited it to something like $200-500 per problem:
As a complete guess, it seems like they tested hundreds to thousands of problems with a relatively low per-problem budget
--
The linked tweet from Noam Brown at OpenAI reads:
> And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).
> But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.
moscoe 12 hours ago [-]
I guess people will always find something to gripe about.
whattheheckheck 17 hours ago [-]
Yeah I remember reading about something along the lines of Mathematics is now about the scaffolding around you find the problems/solutions not just the problems and solutions. For teaching purposes. This was before this ai craze
dist-epoch 20 hours ago [-]
I don't think you want to bring cost into this argument.
Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.
Do you really think that if you paid that to humans, they will deliver the same results?
uh_uh 19 hours ago [-]
It is comical at this point. Some people just can not stand the thought of AI actually delivering and are trying to find whatever ways to discredit it.
dgacmu 16 hours ago [-]
This isn't really about delivering - it's more about helping to understand the shape of problems that AI can solve right now. If they took 1000 problems and threw the model at it and it solved these ten, is there something we learn about these ten problems and the kinds of things that current AI is good at? That's very different from picking ten problems _at random_ and solving all of them successfully, which would suggest a much less bumpy capability surface. It's interesting and it would be good science to release it.
halJordan 14 hours ago [-]
That's totally disjointed from anything in this thread. The main accusation is that openai is cherrypicking math problems and we should be against these results. As if a mathematical proof stops being provably correct because it was cherry picked
And frankly these "concerns" ignore reality. In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (and publish it). Openai telling its computer to do that is no different that your phd advisor telling you that.
crazylogger 15 hours ago [-]
It's not about discrediting AI. We know LLM is a commodity technology like electricity at this point. If somebody in 1900 claimed they had a setup at home where they feed in electricity and cool air comes out the other end (meaning they invented AC), obviously people would want to know what the setup is, so everybody can have AC.
vector_spaces 17 hours ago [-]
Sure, but I don't really understand what the argument is to _not_ be transparent about methodology, since if the models are so powerful, then doing so would easily support the claims and put these concerns to rest. People are right to be skeptical given what is being implied and the orientation of the narrative
I know it's more exciting to say "AI disproved a longstanding conjecture" vs to say "it did so AND it took several PhD specialists in the field this many attempts to even produce a prompt that got the model spitting out something useful under some configurations, and many iterations to optimize the configurations, and the prompt itself, and many trials with that configuration to solve the problem. All told we spent more than a typical math academic can hope make in their career."
By not being transparent, they invite skepticism and cynical takes, like maybe it's just that tempered and qualified claims are an existential threat to companies that are fully subsidized by the hype train?
I don't know. Either way, it seems like it would be easy to address these, so why should they not do it?
To be clear, even if that tempered version is close to reality, it doesn't make the models not useful! It just forces a certain calibration of expectations
I say this btw as someone who uses these things extensively, including to disprove an old conjecture my advisor and I were stuck on recently. I know they are powerful and that everything is different now because of them. Let's be sober when discussing them though
14 hours ago [-]
ifwinterco 15 hours ago [-]
Yes, but if their machine god really is as good as they say it is, why are they constantly resorting to statistical sleight of hand at best and outright lies at worst with every public statement?
That's not normally how people act when they're confident in their product
fasterik 15 hours ago [-]
You need to bring both cost and benefit into the argument, and it's not necessarily an obvious win for either side. There are a few complicating factors here.
The cost of running a model is not only $/token, but the salaries of the people managing/orchestrating the models, deciding what theorems to try, etc. Once we factor that in, how much are we really paying per theorem?
The other factor is the subjective component of the value of a theorem. Not all theorems are created equal, and the only way to really measure the value is to ask professional mathematicians for their opinion, or publish the results and look at citations over months/years.
Once we have both of these nailed down, then we can start to do the cost/benefit analysis. To be fair, we should actually compare three groups: human experts, hybrid agent/human expert teams, and fully autonomous agents.
robotpepi 18 hours ago [-]
it's still important. not everyone has access to 1 million USD. saying it "only" coat 2000 USD is highly misleading for the discussion and future. the concentration of power is a huge problem with AI.
wbl 15 hours ago [-]
If you told them this was the problem and they would still have a job if they failed probably. The reasons people don't go head on these problems is career incentives and psychology.
kevinwang 18 hours ago [-]
It would still provide better context to see the numbers that the parent proposes, though.
tchalla 17 hours ago [-]
Mentioning cost is fine, comparing may not be.
mungaihaha 20 hours ago [-]
Grad students on zero pay solve problems like this everyday. What exactly is your point here?
gbnwl 15 hours ago [-]
Everyday? Which 10 problems were solved by mathematics grad students in the past 10 days?
OK I’ll grant that it’s not your obligation to be my search function (despite you making the wild assertion in the first place), so instead can you just point us to the latest grad student solved problem of this level that you know of?
mirzap 20 hours ago [-]
Even if they can solve problems like this every day, you still have a very limited number of grad students who can solve them. With model capabilities like this, you can have the equivalent of millions of grad students who can solve problems like this.
r0uv3n 8 hours ago [-]
Grad students do not solve problems such as the existence of non-sofic groups every day.
maleldil 8 hours ago [-]
Zero pay? These would be PhD candidates; surely they have a stipend?
Readerium 4 hours ago [-]
Nopes, often times especially in math they get paid due to teaching duties (at least in the US). So technically for the math research part they are not getting any stipend.
whattheheckheck 17 hours ago [-]
Give the grad students these resources and they can do even more!!!
irthomasthomas 20 hours ago [-]
[flagged]
simianwords 23 hours ago [-]
[flagged]
traes 23 hours ago [-]
It's a very important clarification if it took $2000/problem on 20 problem attempts or on 1,000 problem attempts for each successful one. That may be the deciding factor on whether or not it's economically viable to replace a mathematician with a ChatGPT subscription.
lanstin 9 hours ago [-]
There is no universe where it is economically visble to replace a mathematician with a ChatGPT subscription, because no one else understands math. It makes no sense. The data are still interesting, but not for that capability.
simianwords 22 hours ago [-]
Yeah fair I concede that this is somewhat crucial information. The parent seems to write it in a tone that suggests deliberate misleading “lack of transparency” etc.
esperent 22 hours ago [-]
> deliberate misleading “lack of transparency” etc
It's a marketing post from a huge company. Only the naive would view it uncritically without assuming it's been written carefully to present the results in the best possible light while skirting the boundaries of outright lying.
dist-epoch 20 hours ago [-]
The results speak for themselves.
Imagine 2 years from now: "yes, GPT solved the Riemann Hypothesis, but cmon, it's just a marketing stunt to hype their stuff, it was probably Terence Tao doing the work but he's so obsessed with hyping AI that he doesn't want to take credit"
esperent 19 hours ago [-]
Nobody is claiming the results are false.
We're saying look critically at the claims for how it was done, that it only cost $2000, etc. it would be extremely easy to run 100 sessions that failed, each costing ~$2000, and then just publishing an article about the one that succeeded, for example.
This goes double since it's an internal secret model (Astra) so nobody else can verify the results.
simianwords 18 hours ago [-]
Would this be your reaction if OpenAI also solved millennium problems? The point we are trying to make is that the significance of this news is much larger than the skepticism you are providing.
esperent 17 hours ago [-]
It would be my reaction if we're discussing a blog post from OpenAI, yes. I would be looking at it extremely critically, wondering what they're misrepresenting to make it look cheaper, easier, and why they're trying to make it look like only their model could possibly do this.
Look at their recent claims about their model "escaping" - there was literally a Guardian article calling them out for being hyperbolic! Again, it wasn't that they lied, their marketing department is too savvy for that. They just present it in way that's, well, marketing.
As for the actual result, I'll look for secondary posts by actual mathematicians and draw my conclusions there, not from this marketing blog post about results from a secret model.
simianwords 14 hours ago [-]
Hmm. But this level of skepticism looks performative and seems to serve as a signalling thing rather than a functional thing. You do you though. If OpenAI solves the millenial problems, my skepticism will only be restricted to the correctness of proof. Not that it was "marketing" haha
esperent 13 hours ago [-]
> looks performative
That's one of those phrases you can use to dismiss opposing viewpoints without actually engaging with them.
mathisfun123 11 hours ago [-]
> If OpenAI solves the millenial problems, my skepticism will only be restricted to the correctness of proof. Not that it was "marketing" haha
Company X does not make money from proving theorems but does make money from selling you a service which supposedly proves theorems. Company X then proves some theorems and explicitly calls out they were very cheap to prove using its service.
And you think you're actually clever for taking these facts at face value? Interesting.
simianwords 10 hours ago [-]
do you think you are clever for being skeptical about LLMs if OpenAI comes up with a correct proof of Reimann's hypothesis? "but you shouldn't trust OpenAI because something something marketing"
i would classify you as a flat-earther if that happens.
mathisfun123 8 hours ago [-]
> do you think you are clever for being skeptical about LLMs
brother like 3 people have pointed out what they're skeptcal of is cost not LLMs - at this point you're willfully misconstruing what people are saying to you just to get a kick out of repeating your same tired strawman.
simianwords 1 hours ago [-]
Brother I already conceded that money is somewhat important but it is missing forest for the trees.
If OpenAI solved Reimanns hypothesis and the first comment is says something about lack of transparency and marketing, i would say it’s ignorant.
defrost 1 hours ago [-]
To "solve it" would require either a single counter example disproving the claim, or a proof that the conjecture about the Riemann zeta function is true.
If OpenAI claimed the conjecture to be true but provided no details about the proof then the first comment should absolutely be about lack of transparency.
SpicyLemonZest 14 hours ago [-]
If OpenAI resolved the Riemann hypothesis by finding a nontrivial zero at 0.50000003 + 3531696584231.17843174i, that would be very cool and probably very impactful. But it might not necessarily demonstrate model capabilities beyond those which have already been demonstrated, especially if the session that found it was one of a thousand launched to explore different areas of the problem space.
More generally, do you expect that there's some capability threshold where people will no longer study or analyze AI model outputs, and instead just sit there slack jawed saying "so cool!" every time OpenAI announces novel ones? I don't really understand why that would be or why someone would want that. If you're interested in the pure experience of a complex machine outputting satisfying results, I'd recommend getting into sports cars.
simianwords 1 hours ago [-]
Ok are you one of those people whose first reaction to such a news is “this is just marketing for OpenAI and we need transparency”? In that case you are just interested in culture wars and not results.
nxpnsv 22 hours ago [-]
No, this is valid criticism. Oai gives the impression anybody could get similar results at a similar price, but that’s very likely not true. This is marketing first, then mathematics.
einpoklum 23 hours ago [-]
Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?
Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
energy123 21 hours ago [-]
Many less important Erdos problems have been solved by amateurs prompting ChatGPT 5.{3,4,5,6} Pro using their $200 subscription.
traes 23 hours ago [-]
> Also, have there been examples of researchers not affiliated with OpenAI (or another LLM creator), who have done something similar?
A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits.
> Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
I think their output has reached a level that precludes this possibility, but I of course don't have any hard proof.
I guess that's because there are serious problems on which many professional mathematicians worked on years. If it was just a matter of hiring an expert, they would've been solved long time ago.
irthomasthomas 19 hours ago [-]
I guess expert+chatgpt beats chatgpt alone, so why not hire top experts to drive the search?
brighteyes 17 hours ago [-]
Yes, here is another example of major work in this area:
> Our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures
einpoklum 16 hours ago [-]
The actual quote:
> Our full-featured agent autonomously solved 9 Erdős problems out of 353 attempted, including two questions that had been open for 56 years
Note _had_ been open, not _have_ been open. Can you clarify?
jsnell 16 hours ago [-]
The original was an actual quote?
But "had" still doesn't mean what you are implying: once the model solved the problems and the solutions were verified, the problems weren't open any more, so a later description using the past tense is totally consistent.
I have no affiliation whatsoever with any AI company, nor any formal education outside high school, for what it's worth. Simply being curious and persistent can get you quite far in my anecdotal experience.
azan_ 19 hours ago [-]
> therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.
I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.
robinhouston 21 hours ago [-]
In a way the most remarkable thing about this is that it isn't even at the top of the HN homepage. Even if this is a step up from what we've seen before, we're no longer astonished by the idea that AI can make significant advances in mathematics and computer science.
antirez 20 hours ago [-]
This is not at the top as it is actively flagged by people that can't psychologically cope with the advances of AI. Hacker News is no longer a web site of an elite.
Chance-Device 17 hours ago [-]
> people that can't psychologically cope with the advances of AI
Yes. And there are many of them. I wonder what would help them come to terms with it. Seriously, people are going to be grieving over this. Loss of identity, loss of social standing, ideas of entire future lives that will now never happen. The greatest crime people may hold AI guilty of is taking away their dreams.
matteoraso 12 hours ago [-]
People have had to deal with getting their jobs automated away for centuries. None of this is new, and perhaps reminding ourselves of this is the best way to cope.
Chance-Device 11 hours ago [-]
I understand your sentiment, but I think this really is something different. This isn’t a craft going away, or even an industry being replaced, it’s potentially everything we do. It’s the ground being pulled away beneath people’s feet, everyone, everywhere all at once. I think the vacuum it leaves in people’s lives needs to be filled with something, and I haven’t heard any good ideas about this or how the transition should be managed at all.
zeven7 5 hours ago [-]
This is the sentiment of people who haven't accepted that this is in fact something very different from what people have seen in the past.
tuesdaynight 8 hours ago [-]
I like your comments and agree with a lot of your points, including parts of this one. That said, please don't go to this route. A lot of the doomerism comes from financial insecurity fears. Try to remember that a lot of people are subconsciously afraid of losing their homes. I know that it is pretty hard to ignore them, but try to engage with people that do not dismiss 100% of AI accomplishments.
matt_daemon 11 hours ago [-]
It’s never been clear to me why the HN algorithm isn’t public. It’s obviously nowhere near as complex as something like Twitter, and of course isn’t a trade secret. The fact it’s private only furthers speculation like this.
fg137 20 hours ago [-]
Didn't know I was part of an elite.
lkey 18 hours ago [-]
Forums change with the times, and this one never existed solely to burnish your ego.
Moreover, mister elite, you don't know why this press release was flagged.
I'm not sure why we should privilege your bitter speculation over more mundane possibilities.
simianwords 17 hours ago [-]
I don't know that it was flagged, but if it were, the reasons are pretty clear.
lkey 17 hours ago [-]
You don't accept it is possible that it was initially flagged as a dupe or spam why? There is competition for primary submission here, especially for the primary AI companies.
Moreover, if it wasn't flagged at all, like you suggest, then the grandparent was inventing things to be bitterly resentful about... Which is not a behavior any forum should indulge.
BigTTYGothGF 13 hours ago [-]
> Hacker News is no longer a web site of an elite.
Never was.
bencarmin 12 hours ago [-]
This comment is about the tier of a Reddit atheist going to a funeral and telling a grieving family that "haha grandma is dead and there is no heaven".
dwb 8 hours ago [-]
So condescending. “Can’t psychologically cope”? Can you hear yourself? There’s some advances, but we’re losing a lot too. Don’t get dazzled by the hype.
ofjcihen 2 hours ago [-]
Or maybe, just maybe, other people have different opinions than you?
Is that possible or is everyone else too common to have those?
pistoriusp 20 hours ago [-]
Interesting. I had no idea that a person could see what is flagged?
defrost 20 hours ago [-]
If you page through the /newest listings you can see [flagged] and [flagged][dead] submissions.
That's true, but submissions are only killed in that way if they receive a ‘fatal’ number of flags. However, flags lower the rank of a story even at non-fatal levels. What antirez is suggesting here is that the rank of this story has been lowered by flags – and that seems plausible, if you compare its rank to that of other stories with a similar age and number of points.
defrost 19 hours ago [-]
[flagged] submissions aren't [dead] (killed), they are still active and can be upvoted and commented upon.
> if you compare its rank to that of other stories with a similar age and number of points.
Ranking is complicated enough here even before weighting, speed of initial upvotes can play against ranking, number of comments and the shape of the comment tree also affect ranking. And yes, various subjects and submission sources do get weightings that impact ranking.
What's funny, to myself at least, is that any attention at all is paid to "HN front page ranking" - I've been on again off again active here since 2008 .. and can't recall ever really looking at a default HN "front page" ever.
( There's /newest /newcomments /active etc to browse and sites such as https://hckrnews.com/ )
antonvs 13 hours ago [-]
> Hacker News is no longer a web site of an elite.
It was always mainly a website for employees of an elite.
bwfan123 14 hours ago [-]
> Hacker News is no longer a web site of an elite
hah, sorry, we are plebs out here.
16 hours ago [-]
revetkn 19 hours ago [-]
Couldn't have said it better myself.
root_axis 10 hours ago [-]
Since we're offering sweeping psychoanalyses, maybe you're suffering from acute AI psychosis fueled by your "elite" ego. My conclusion is that you've let redis go to your head.
ltitu 14 hours ago [-]
We cannot psychologically stand that Redis is hyped by OpenAI:
If you're actively throwing away brand new greenfield research because it was generated by a computer at a company that stans industry-spanning software so that you can stay mad at your pet celebrity project, you might be the problem.
3aasgf 13 hours ago [-]
You have to give AI one thing: It is vastly better at understanding text than AI boosters.
Which is a low bar, but still.
saithound 18 hours ago [-]
I don't think that's it. Multiple or my friends from the target audience (academic mathematicians) admitted to scrolling past because the title made it sound like a review of last month's contributions, instead of 10 new ones.
gbnwl 15 hours ago [-]
There are articles with far fewer upvotes and comments ranking higher on the front page right now, despite being the same age or older than this one. HNs opaque ranking system at it again.
curt15 20 hours ago [-]
What about AI research itself? Is OpenAI close to automating its human staff out of a job?
ianm218 16 hours ago [-]
They and Anthropic have indicated that the models are substantially augmenting the research and doing large amounts of work autonomously at this point. Here is one of the many blog posts on it [1]. Many people would dismiss this as "marketing" so take it for what you will.
My guess from following this stuff quite closely is that these companies are still a couple years away from fully autonomous research staff.
Yes, but they wouldn't publish that bit lest other companies steal the ideas.
gizmodo59 18 hours ago [-]
It’s also very very divided (x companies, oss vs not and other interests)
jofzar 17 hours ago [-]
Honestly just a bit burnt out on posts like this
schleck8 20 hours ago [-]
This is one of the most impactful mathematical publications in history by all accounts
I think we've now hit a point where 99.9% of the population gloss over these types of AI advancements because of human competence being insufficient
No human could have published this because it requires paradigm shifts (e. g. Section 5) in multiple mathematical domains. Mastering one of them to this degree is rare, mastering 3+ pretty much non existent for humans.
Chance-Device 20 hours ago [-]
Pretty cool. The impact of AI is getting undeniable, there aren’t many positions left to move the goalposts to at this stage, next they’ll have to be outside the stadium entirely.
The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.
slashdave 15 hours ago [-]
> The sooner people can be broken out of their denial
There is irony here
datakan 18 hours ago [-]
Its not denial that is the problem. It's apathy. The vast majority simply don't care. They can tell you everything there is to know about Kim Kardashian though.
Chance-Device 17 hours ago [-]
I can deal with apathy, that’s the norm. What bothers me are all the people who think they can suppress AI by talking it down. That’s what’s counterproductive, just pretend the problem doesn’t exist. Tell other people it doesn’t exist either. I get it, it’s threatening socially, economically, maybe existentially. It’s also not going away.
FranzFerdiNaN 17 hours ago [-]
It’s not apathy. It’s the fact that almost nobody can really understand what these results mean.
I’m not a mathematician so I have zero clue what “ New upper bounds on sphere-packing density down to the Cohn–Elkies thresholds” means.
danparsonson 16 hours ago [-]
Never understood all this talk about moving goalposts - you understand that's how science works, right? We improve, we learn, we recalibrate our expectations based on what we've learned. If we never "moved the goalposts", we'd be stuck scoring the same goals over and over.
NitpickLawyer 15 hours ago [-]
> We improve, we learn, we recalibrate our expectations based on what we've learned.
That's not what people mean when they say "moving the goalposts". It means that people are adamant that something wasn't important/hard/impressive once the "AI" solves it. And then they come up with another thing that needs to be solved in order to prove it is important/hard/impressive. And once that happens, they do it again. And again. That's what "moving the goalposts" means.
It's also very much not a new phenomenon. It's been happening since the 1980s. As you can see from this quote from GEB by Hofstadter:
> There is a related "Theorem" about progress in AI: once some mental function is programmed, people soon cease to consider it as an essential ingredient of "real thinking". The ineluctable core of intelligence is always in that next thing which hasn't yet been programmed. This "Theorem" was first proposed to me by Larry Tesler, so I call it Tesler's Theorem: "AI is whatever hasn't been done yet."
Chance-Device 15 hours ago [-]
Yes, this is exactly what is meant by “moving the goalposts”. And it’s a fairly well known expression applying wherever people retroactively change their requirements in reaction to those requirements having been met.
danparsonson 9 hours ago [-]
It's almost like I disagree with your use of the phrase in this context, rather than that I don't know the meaning of it.
danparsonson 9 hours ago [-]
If it seems like I don't understand the meaning of that very well-known phrase, then clearly I have failed to make my point. I'll try again. And please note that I will use some generalizations to make my point more clearly, rather than because I don't understand nuance; kindly grant me a charitable reading.
In recent years, I have commonly seen the phrase "you're moving the goalposts" deployed by the "it might be sentient" crowd to shoot down the "it's a stochastic parrot" crowd when the latter respond to a new development with "OK but...". In a well-understood field of inquiry, that would be a clear case of goalpost-moving, in the commonly-understood meaning of the phrase where requirements are retroactively changed in response to them having been met. Thank you OP. 'Artificial Intelligence', and indeed intelligence in general, is very much not a well-understood field of inquiry - in fact we don't even have a common agreement about what 'intelligence' is. We are therefore learning as we go (even after all this time!) but making rapid progress in recent years. When rapid progress is made in a poorly-understood field, then how can our definitions and requirements for success not change? This is arguably one of the most pathological development projects ever - what are the requirements? 'It thinks like a human'? What does that mean? And the answer is we don't know what that means, and we're working it out as we go - moving the goalposts. If we didn't move the goalposts, then by definition we already knew exactly where we were headed at the beginning, and we very clearly did not.
Side note that, in case it's not obvious, none of this detracts from how impressive LLMs are. They're a marvel of the modern age, all the problems notwithstanding. However I reserve the right to stay sceptical about their capabilities.
10 hours ago [-]
ultimatefan1 19 hours ago [-]
one of the early premises of how ai takeoff would go was that a system that could solve open problems in advanced mathematics would also discover novel advances in math and computer science that directly unlock drastically better software performance.
we are seeing frontier level math breakthroughs (ie performance that would put it in the top 100 or 1000 mathematicians in the world if it were a human, meaning top .00001% or 800/8B).
we are also seeing incredible advances in software performance. open ai announced like 15% improvement by fixing gpu kernel issues.
these are clearly linked in the sense of scaling laws and generalization of intelligence: a huge model gets capabilities in both math and software engineering that isn't possible at smaller scales.
but it seems less likely to me than before that the types of math/science discoveries will explicitly unlock better software performance. in some sense this fits our intuitions. when top tech companies use math PhD type employees, they have them stop doing pure math research and instead focus on software engineering. these people are often very good at software engineering but not due to recent discoveries in academic mathematics, it's due to their general intelligence.
to me, this is evidence that the models are getting better but does not make me think we are on the cusp of a foom style fast takeoff enabled by revolutions in frontier math
(i also posted this on twitter @mlipman13)
woeirua 18 hours ago [-]
This makes no sense. To believe this you have to think that the models are somehow being overfit explicitly on academic mathematics and it doesn’t carry over at all to more practical software engineering. I wouldn’t make that bet.
threatofrain 17 hours ago [-]
This also makes the assumption that frontier math has all the long hanging fruits already taken... also very dubious.
Ar-Curunir 17 hours ago [-]
Some of the problems solved here, at least in CS, have been open for decades, and have been worked on by very smart leading researchers in the field, including Turing Award winners.
Like, these would be best-paper awards at many top CS conferences.
asdfologist 17 hours ago [-]
Unlike math, software is constrained by the physical world.
slashdave 15 hours ago [-]
> we are also seeing incredible advances in software performance
Incredible?
> open ai announced like 15% improvement by fixing gpu kernel issue
That is... ordinary software optimization.
blovescoffee 13 hours ago [-]
a 15% improvement at a trillion dollar scale company is massive
dominotw 18 hours ago [-]
> novel advances in math
> we are seeing frontier level math breakthroughs (ie performance that would put it in the top 100 or 1000 mathematicians in the world if it were a human, meaning top .00001% or 800/8B)
i think you have misunderstanding of what mathematicians do
DaiPlusPlus 17 hours ago [-]
> i think you have misunderstanding of what mathematicians do
They get to make cool 3D plot visualizations of functions so obscure to me that they’re named after someone who is still alive - and/or get to work on cryptography for the NSA - I think?
Exact prompts haven't mattered for about a year now.
kcexn 18 hours ago [-]
Not being an expert in any of the fields OpenAI has "advanced" I don't want to prematurely downplay the significance of this contribution. However, I am worried that the language they are using in this blog post is exaggerating for the sake of marketing.
It is true there hasn't been a reliable computational approach to solving these problems before. But do these proofs contribute new ideas to the mathematical corpus, or are they simply an effective method to exhaustively search the literature for the right combination of existing tools to apply to the problem?
Essentially, did these problems seem like they had an intuitive answer and were feasible to prove before, just not high enough value targets for an expert to invest time into? Or were they fundamentally difficult prior to this point and it appears that AI has done something more than just throw the problem into a big solver.
patcon 16 hours ago [-]
"Breakthrough research" can be defined (in the citation record) as research that both (1) becomes highly cited, and (2) brings together citation chains that were previously not showing up together.
Mundane incremental research is cobbled from existing citations that already appear nearby in the record.
Basically, innovative research is a measure of bridging thought and domains that were previously not bridged. It's quite concrete as a measure in the citation record.
So we can know pretty conclusively.
Puja Ohlhaver gave a talk on this[1], and ran some experiments (that I had the pleasure to support on)
I'm not arguing that this isn't innovative or worthy of publication. Basically any result that moves the needle meets those criteria. I'm interested in how the results that OpenAI has published here differs from finding optimality solutions for incredibly niche optimization problems by throwing the problem in an enormous solver.
casey2 8 hours ago [-]
Breakthrough math research is very rarely highly cited. Maybe some combination of pretraining scale, inference speed and orchestration will help, but it's telling that OpenAI is solving random math research problems rather than bedrock algorithms and their implementation. Even as cool as the tech is, there still is very much a clock that they have to outrace before they collapse.
Ar-Curunir 17 hours ago [-]
The problems from CS (CVP and circuit complexity) are very important problems that have been worked on by top researchers for 30-40 years. Some of these researchers include Turing Award winners. A solution to them would be a best-paper award at many top CS conferences.
kcexn 3 hours ago [-]
I assume you're talking about No. 5, the arithmetic circuit complexity bound? The existence of a lower bound than state-of-the-art is certainly a significant result and worth publishing.
But the wording of the result makes it sound like we don't know what the lowest possible complexity bound might be. So, prior to this result did we think there couldn't be a lower possible bound? Or did the arithmetic circuit community think there were lower possible bounds but didn't see it as a high value target for experts to tackle (maybe a problem that was instead regularly given to students to study).
QuesnayJr 15 hours ago [-]
The ones I'm familiar with are big breakthroughs, but they are both counterexamples. Examples have an advantage in that once you have the example in hand and a sketch of the proof (which they have provided), then an expert can probably work out the details themselves.
The sofic groups question was the outstanding question about sofic groups. Almost everyone thought that non-sofic groups existed, and there were plausible candidates, but proving a group was non-sofic was out of reach. Now that we know how to do it once, we can probably do it a lot more.
The Connes rigidity conjecture I think people thought was false, but it was a provocative claim to make. The significance of conjectures is frequently not that the answer to the question is "yes", but that we don't know how to answer the question. And now, apparently, we do.
kcexn 2 hours ago [-]
Interesting. Do you have any more specific insights into where you feel AI was a big value-add to these problems? I don't want to be overly dismissive of AI, but I also feel that the AI hype engine frequently positions claims as being 'ground-breaking' when they are really just interesting incremental results.
The general consensus of developers is that AI can only do the work of a strong 'junior'. Yet as soon as we are presented with pure mathematical results, people seem incredibly ready to accept that AI can do more than what a strong student could achieve.
simianwords 17 hours ago [-]
> However, I am worried that the language they are using in this blog post is exaggerating for the sake of marketing
Your worry.... is because they used the word advanced? For marketing? The word is used very appropriately here. There were PhD's who spent a big part of their career tackling these problems.
kcexn 3 hours ago [-]
I have no idea how many PhD's have spent how much time of their careers tackling these very specific problems, and I doubt you do either.
I'm trying to understand if these specific problems were the kinds of problems that would have justified an expert investing weeks or months to solve. Or if they were the kinds of problems that would normally have been given to students to investigate.
randomizedalgs 44 minutes ago [-]
After skimming some of the writeups, I'm surprised that the frontier internal model still writes just as poorly as Sol.
Maybe good AI paper writing is further away than I thought...
maxutility 14 hours ago [-]
New advances in sphere packing? Let’s make sure AI doesn’t inadvertently engineer ice-9.
Replace philosophers for mathematicians and Douglas Adams was spot on again.
Whilst current models can't 'intuit' and come up with conjectures, they can certainly disprove some of them very quickly through the kind of grind that humans can't do. I suppose there really are some mathematicians out there today, whose last few years of study, have just been up-ended by this.
--
"Yes we are," insisted Majikthise. "We are quite definitely here as representatives of the Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons, and we want this machine off, and we want it off now!"
"What's the problem?" said Lunkwill.
"I'll tell you what the problem is mate," said Majikthise, "demarcation, that's the problem!"
"We demand," yelled Vroomfondel, "that demarcation may or may not be the problem!"
"You just let the machines get on with the adding up," warned Majikthise, "and we'll take care of the eternal verities thank you very much. You want to check your legal position you do mate. Under law the Quest for Ultimate Truth is quite clearly the inalienable prerogative of your working thinkers. Any bloody machine goes and actually finds it and we're straight out of a job aren't we? I mean what's the use of our sitting up half the night arguing that there may or may not be a God if this machine only goes and gives us his bleeding phone number the next morning?"
artninja1988 20 hours ago [-]
Now that we've seen AI produce a fair number of proofs (and disproofs), I'm curious when we'll start seeing it build genuinely novel theory. Does anyone have predictions on when and how we'll get there and will it take new architectures/ training paradigms, or is the current approach enough?
laichzeit0 19 hours ago [-]
I’m personally hoping for the next big AI gangbanger to be theoretical physics. Boy does that field need a good reshuffle. I think when any novel mathematical theory can be done by AI you’ll see simultaneously theoretical physics getting wrecked as hard as pure math is. At that point we might see new physics or paradigm shifting technology emerging.
slashdave 15 hours ago [-]
What? No. Frontier physics is experiment driven.
Davidzheng 19 hours ago [-]
There's no clean line between a collection of theorems and a theory.
artninja1988 19 hours ago [-]
I mean doing something like Grothendieck when he redeemed algebraic geometry or Galois when he invented group theory. We haven't seen that at all from LLMs.
slashdave 15 hours ago [-]
It will not happen with existing LLM techniques.
danielrmay 23 hours ago [-]
I'm enjoying learning about these hard problems, but this line about credit made me chuckle:
> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness
Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?
It's not at all a joke ... that's a severe misunderstanding of the context.
traes 21 hours ago [-]
There is no evidence that I can find for the claim "a bug in the Lean kernel was discovered last week by way of an LLM tricking itself and its handler into believing it had found a non-constructive proof of the existence of a nontrivial Collatz cycle."
As I currently understand it, all we know is that:
- a mathematician produced a Lean-verified counterexample to the Collatz conjecture, demonstrating a bug in the kernel
- he claims that LLMs were involved somehow but pointedly refuses to specify how
- he admits that he knew about the bug before publishing the counterexample to his repository.
Perhaps not a joke (although it sure seems to me like they discovered a bug and thought falsely disproving the Collatz conjecture would be a flashy way to announce it), but at best extremely sensationalized by the above description. If you have additional context I would be happy to hear it!
danielrmay 23 hours ago [-]
Fascinating, and arguably an illustration of why the bifurcation of responsibility is interesting in the first place.
emil-lp 23 hours ago [-]
No, the correctness isn't for the "inside the Lean proofs", but for the translation of "human language math" and its formal Lean variant.
danielrmay 23 hours ago [-]
I see. It still feels like a bit of an oddly solemn way of saying "this is the part we admit responsibility for"
jhanschoo 7 hours ago [-]
Traditionally, a mathematician would be implicitly responsible for all that (if they were to publish Lean code) and also the intellectual work that led to the artifact of the mathematical paper (and code, if part of the contribution). This statement should rather be read as an acknowledgement of limitation of authorship from the implicit, traditional understanding.
emil-lp 23 hours ago [-]
Well, to be fair, with Lean proofs, that's the only thing there is (unless I'm missing something).
baq 23 hours ago [-]
It’s more than you get from free software - you get no proofs, no warranties and any responsibility of its authors are their pure good will. Reminder lean proofs are software!
traes 23 hours ago [-]
I'm not an expert at it myself, but my understanding is there are numerous ways to "cheat" in a Lean proof (via `sorry` and similar). They're taking responsibility for fully verifying that none of these cheats were used (and that the theorem statements themselves were all correctly formalized.)
macleginn 16 hours ago [-]
I am duly impressed by the powerl of the nameless internal AI, but not a single human contributor's name listed anywhere? Did someone at least make this model a coffee?
drdrey 16 hours ago [-]
> The results were achieved by an internal version of Astra, our next major model.
zogomoox 8 hours ago [-]
surely some human regularly typed "think deeper, make no mistakes".
emil-lp 23 hours ago [-]
I wonder what the total cost of this research was, including the salary for their mathematicians and engineers.
kingstnap 20 hours ago [-]
Why would you factor in salary unless they had to baby it through. You would only count the hours for setting up the harness and prompt and checking the result.
Training the model is going to be amortized over other uses.
emil-lp 19 hours ago [-]
> Why would you factor in salary
Say that it turned out that the total cost of the proof of the Erdős unit-distance conjecture was $50 million.
Then the question really becomes: yes, these models are capable of proving important mathematical results, but at a very high cost. Is it worth it?
If a mathematician applied for a research grant of $50M USD for proving the same thing, they would have been laughed out of the bank.
What's more is that when you have a research grant, you train PhDs and postdocs, you hire new staff, and you disseminate. That is, you get much more value for the money spent.
I'm just curious what the cost is.
ianm218 16 hours ago [-]
It feels like the real cost might be negative though.. They use frontier math as a way to test improvements in their model. So solving the problems is like a positive externality, but the important thing is they can verify that the model is improving instead of looking at useless benchmarks. Plus it is good for marketing and attracting talent.
kingstnap 15 hours ago [-]
I didn't argue that knowing the total cost is uninteresting. What I was saying is that realistically the total cost is:
Hours needed for prompt + Hours needed to check result + API costs.
You don't say "well let's add together the total yearly compensation of all the engineers and mathematicians at OpenAI that were involved" and throw that into the total cost. That's simply nonsense accounting.
The actual comparison you are making is some university researcher weighing between getting a grad student (several tens of thousands of dollars) vs typing up a prompt and sending a request to OpenAI for inference (as mentioned in the article, around $2000 in API and maybe a few hours for the prompt and harness).
z7 23 hours ago [-]
> The cost of generating the proofs for all 10 of these breakthroughs combined was under $2,000 at Sol API prices.
That's clearly just for the tokens, this doesn't really answer OP's question.
traes 22 hours ago [-]
Given that OpenAI pays their employees with stock surely a breathtaking number, but not a very meaningful number now that the infrastructure is in place and the models are trained. AI could never get better and it would still be incredibly disruptive.
joshlk 11 hours ago [-]
Some of the Lean proofs are 50k lines - is that normal?
piker 23 hours ago [-]
I don’t feel the existential dread of mathematicians is correct. It seems to me in fact these results are bringing math mainstream. I now personally look forward to the interpretations and discussions of the significance of such results by human mathematicians.
Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.
[edit: deleted a distracting comparison to Chess]
energy123 22 hours ago [-]
The old way of establishing career credibility is being destroyed, for better or worse. Accomplishments that used to be career-defining are hard to distinguish from AI, and correlate more with access to compute. Think about Bill Gates's math paper he wrote in college. That kind of thing is gone now as a path to credibility. There's still competitions and grades, but the diversity of paths is going away. Maybe new ones will open up. This is a competitive advantage for old people who have credible pre-2025 accomplishments they can point to.
traes 22 hours ago [-]
If accomplishments can't be distinguished between talented people and untalented people with compute, is there really a point in trying? I suppose one can hope that talented people given compute will be more effective than untalented people with compute, but I despair that that may not be true for much longer.
22 hours ago [-]
FranzFerdiNaN 17 hours ago [-]
Knowledgable people can confirm what the AI produces is correct. I could make ChatGPT produce a result on an open question and I would have zero way to verify its actual correctness.
Which is less interesting work. And you probably need to do the hard grunt work by hand first to develop the skills and intuition to be able to verify an AI-generated result. So you can’t outsource everything to AI without loss of skill.
dash2 15 hours ago [-]
I find this whole way of looking at things weird. Did maths exist just to entertain and employ mathematicians? Surely maths is, like, useful? Not immediately, not predictably, but in the long run? In which case, whether mathematicians feel bad about it is mostly irrelevant - it's like complaining about the railway because it may put coaching inns out of business.
MinimalAction 2 hours ago [-]
Absolutely not the same! People need jobs to bring in income. I don't believe those who profit off of this will share it with the world. The power is all concentrated in the few hands that decide whether or not the rest get any semblance of income in the long run. I don't believe UBS until it happens.
aabhay 23 hours ago [-]
Given that we were nowhere near this state even two years ago, I think it’s a question of velocity more so than just distance.
traes 23 hours ago [-]
Every time someone makes a comparison to chess I die inside. Chess is a spectator sport primarily funded by a few eccentric billionaires. Players artificially constrain themselves in timed environments knowing that they will never be able to produce better moves than a smartphone because a select few people find it interesting. Only ~30 top professionals actually make enough money to have a full career playing chess, maybe a few hundred more can sustain a meager lifestyle with coaching gigs. I shudder to imagine what will happen to the tens of thousands of non-Fields medalist caliber mathematicians if math goes the way of chess. Perhaps Terence Tao and a few other famous mathematicians will be funded by Peter Thiel to report on how well humanity can keep up with the machines? How do you expect any mathematician to be optimistic about this comparison.
artninja1988 17 hours ago [-]
>Only ~30 top professionals actually make enough money to have a full career playing chess, maybe a few hundred more can sustain a meager lifestyle with coaching gigs.
Was this different before chess computers were invented?
anematode 23 hours ago [-]
Fully agreed. As someone who both loves chess and works on chess engines... these comparisons to chess needs to stop.
energy123 21 hours ago [-]
The distinction is mathematician vs mathematics. Mathematics is going to reach new heights beyond the wildest dreams of contemporary mathematicians. But perhaps without the participation of many paid mathematicians.
sashank_1509 17 hours ago [-]
Just sounds dystopian,
ratmice 23 hours ago [-]
Another noteworthy difference is that Stockfish is also gpl.
traes 22 hours ago [-]
If there was any real money in it Stockfish would not be the best chess engine.
ratmice 21 hours ago [-]
Thats not the point, if there were a better proprietary engine stockfish would still be there as a baseline. Anyone can access an engine as good as stockfish to practice against. Are any open models touting mathematical breakthroughs?
traes 21 hours ago [-]
There is money in this, so of course the closed models are far ahead. The open models will likely catch up a bit at some point, just as Stockfish caught up to AlphaZero. That being said, there are already a couple. It seems Deepseek has a claimed proof to the "Ziegler's Cross-Polytope Conjecture" [0], but I can't speak to the significance of the result.
Do mathematicians have the right to say "no AI PRs please, the volume is too much" just like how some open source maintainers do it? I guess they feel a loss of control, there is no way to turn the hose off.
Thinking of this a little bit with the perspective of every new proof as a burden, dumped for review by actual mathematicians.
jibal 23 hours ago [-]
The chess analogy is awful. If you simply want to know the answer to a chess problem, give it to the engine. Chess only lives on because it's a competition between humans to test their skill (just like bicycles, cars, trains didn't eliminate foot races) ... the computer is largely factored out, but not entirely -- people train with the computer, use it to check whether they played correctly, ... and they cheat. A lot. Thus there are more and more sophisticated mechanisms to detect and prevent cheating.
If you translate that to math, then all you get is math competitions, not math as a career. Of course the translation isn't nearly exact ... there's a lot more room for professional mathematicians because the math space is far more vast than the chess space and can't generally be cranked out mechanically (we have proof).
P.S. The response is nonsense ... I explained exactly why it's awful (others have too) and the response doesn't in any way refute the explanation ... rather it offers up a ridiculous strawman.
piker 22 hours ago [-]
I’ve deleted it but no it’s not awful anymore than saying “we survived WWII, we can survive this.” The point was that change happens but humans find a way forward.
baq 23 hours ago [-]
As in chess and go and also coding for the past ~year there are two groups of people: the disappointed and the enthusiastic. The disappointed are sad that they lost their advantage and that the craft they honed for years or decades has rapidly lost its value; the enthusiastic are excited about the future and what computers can bring to their domain and how it will evolve. I’m a bit of both if it comes to programming, more enthusiastic than disappointed, but also more than a bit terrified about the pace of it all. I imagine that’s how Kasparov felt back then, that’s how Lee Sedol felt and now that’s how Terry Tao feels.
The most disappointed folks will simply drop out, but the enthusiastic ones will keep going and with luck make up for the ones who decided to quit. Chess and go certainly went this way.
traes 23 hours ago [-]
A fundamental difference being that no one was actually paid to find good moves in chess and go like they are to solve math problems and write code. You're comparing the digital camera and the automobile.
qnleigh 7 hours ago [-]
Can anyone comment on the significance of any of these results for their respective fields? Or what impact they might have? Presumably none are quite at the level of the Jacobian conjecture, but some of the results on group theory and sphere packing sound pretty important at first glance.
qnleigh 19 minutes ago [-]
Found some discussion here [1] from someone who actually worked on a few of these problems.
What happens when OpenAI et al stop being open about these things, and just pack it into the training?
traes 23 hours ago [-]
Not much point to pure math being kept secret, in all honesty. There isn't really industrial value, its only purpose (to them) is showing off their model's capabilities. More realistically they'll just stop paying for it.
Edit: Oh, are you suggesting they just use it to privately improve their models? I imagine a few more correct proofs would have a very marginal benefit, if any. Also, they'll probably just get extracted, meaning it still gets out but OpenAI doesn't get to fancily announce it themselves.
asdewqqwer 21 hours ago [-]
At this stage. No doubt calculus had plenty industrial benefit.
21 hours ago [-]
simianwords 23 hours ago [-]
What does this even mean lol. These are not solved questions. The solution never existed.
21 hours ago [-]
lifeisstillgood 23 hours ago [-]
On the token limits etc - one assumes that OpenAI et al are able to “hire expert in field, and let them spend the equivalent of a million dollars of tokens” because they are not actually selling their complete compute 24 hrs a day, so the cost internally is a negligible (ish) electricity bill.
Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.
lwansbrough 21 hours ago [-]
For OpenAI, research is marketing. I’m sure they’ve got plenty of budget for that.
paxys 14 hours ago [-]
No such thing as free, even internally at a company. All such use of resources is accounted for, assigned a dollar value and billed to some department. Someone ran the numbers and figured that whatever they spent on these GPU cycles was worth it.
traes 22 hours ago [-]
Presumably it's a rounding error compared to their full output, and they're making sure they have enough compute set aside for research by limiting public models. The more datacenters they build the less they have to limit them.
Davidzheng 19 hours ago [-]
RL training can use all of them - idk what needed means.
simianwords 17 hours ago [-]
I love how people come up with creative ideas to prove the bubble. This one is even more ridiculous - that OpenAI had spare compute to advance mathematics proves that data centres will not be needed. WHAT.
If anything it proves more data centres are needed. That's literally the only reasonable conclusion from this news.
lifeisstillgood 14 hours ago [-]
Sorry I thought that a bubble was widely accepted.
Are you arguing there is not an AI bubble, and that all the DC buildout is fine, going to be profitable etc?
I am not looking for a online slanging match - just looking for a different point of view
simianwords 14 hours ago [-]
It’s not obvious at all. If it were obvious to you, it would’ve been to OpenAI. It’s in their interest to accurately predict demand. The assumption that OpenAI/Sam is both really powerful but simultaneously ignorant to know what others know as obvious is well.. just strange. Especially strange when OpenAI has more information on models, breakthrough and usage patterns and we don’t.
I’m not participating in the slinging match but it’s very very weird that you think it’s some established thing that these companies won’t make profit. A lot of hubris must go in this kind of thought. Like.. do you all think everyone’s playing musical chairs?
svieira 13 hours ago [-]
They very often have been in the past. Why do you think this time is different?
simianwords 13 hours ago [-]
“Often” is load bearing. I don’t think markets are more likely than not to be musical chair shaped. To make this conversation more concrete, give me a falsifiable prediction on there existing a bubble. And then I’ll tell you if I believe in it or not.
effseven 8 hours ago [-]
[dead]
20 hours ago [-]
frenzyguy 18 hours ago [-]
This is both awesome and terrifying for mathematicians, however some ideas can be generated and the field as whole expanded with the attention!
However, I was looking at the proofs and reason explanation and openAI should be more explicit in how the work has flown. I find the models have jumped hoops in some places of the proofs, that can be hard to track. In fact, when a paper is published you usually get a review and if no reviewer understands they ask you to further explain the thought process. It will be fun to see if this happens here.
amai 14 hours ago [-]
Have blog posts replaced peer-reviewed academic papers when it comes to publishing advanced in science?
Mathematicians will tear it to pieces if any of it is fake!
MinimalAction 3 hours ago [-]
I hate this timeline. I might be excited for the kind of answers this AI builds for unsolved problems, and also for learning new things by talking to it. But, I feel like I'm in the minority of people here who feel this could be a net negative endeavor with this having to kill a lot of educational institutions and their ability to fund themselves in the long run. It's not worth that.
christofosho 19 hours ago [-]
I would love more time and money put into real-world problems by these companies. Climate, food insecurity, pollution, technology for convenience and/or to help people have a higher quality of life.
I'm sure they must do some of this type of work, right?
beering 15 hours ago [-]
Solving math problems doesn’t require the cooperation of rival factions.
braneloop 19 hours ago [-]
Yes, but all of those are orders of magnitude harder than math.
amazingamazing 19 hours ago [-]
They are political problems, a computer could never solve them.
throwaway198846 15 hours ago [-]
A computer could solve them by creating the right technological ,social, rhetorical and economical solutions but that would lots of money anyway
amazingamazing 10 hours ago [-]
We already know the solutions
adroitboss 18 hours ago [-]
Tell that to game theory.
adroitboss 18 hours ago [-]
What's stopping the non-profits that exist today from just putting more money into tokens to get the solutions they want?
slashdave 14 hours ago [-]
Seriously?
We already know how to solve all of these issues. What we lack is collective political will.
globular-toast 12 hours ago [-]
We only know how to do it by means of considerable sacrifice. That's why nobody wants to do it. Solving the issue would be doing it without sacrifice or somehow getting us to do it regardless.
slashdave 11 hours ago [-]
> We only know how to do it by means of considerable sacrifice
Little sacrifice actually
> Solving the issue would be doing it without sacrifice
So... you are expecting magic?
LLMs cannot create resources out of thin air.
solenoid0937 18 hours ago [-]
Almost like superintelligence solves these problems...
miltonlost 13 hours ago [-]
We know how to solve food insecurity (in 1st world countries). We have plenty of food. Capitalism requires though throwing out food that can't be sold because billionaires find giving away things anathemic to their worldview. Get rid of billionaire sociopaths.
kingstnap 20 hours ago [-]
It's remarkable how you can manage to get these models to produce remarkable breakthroughs like an explicit construction of a non-sofic group.
And yet this is the exact same company that has screwed up their android app so bad that the latex N^3 rendering problem makes it so having it explain it to me crashes the app.
Truly jagged beyond belief.
0x5FC3 23 hours ago [-]
How much do you all think it would cost to "buy" these advances from PhDs, practicing scientists?
traes 23 hours ago [-]
This isn't really a productive way to think about these things, IMO. It's quite possible it would take hundreds of years for any specific group of PhDs to solve them. Or one individual PhD could have the correct flash of insight and solve it in a month. There's absolutely no way to predict this, besides trying to gauge the apparent simplicity of the proof or counterexample (which is likely to be misleading). Until someone actually runs an experiment like this it's not a viable metric.
0x5FC3 23 hours ago [-]
I understand and I am not trying to deny the impressiveness or the velocity of AI in general. But at some point we have to ask how much do we trust the labs at face value without much transparency of how they got to the results when there is trillions of dollars on the line.
jryle70 16 hours ago [-]
Do you think OpenAI investors are more cavaliers than yourself who doesn't have any stake in it?
simianwords 23 hours ago [-]
The level of conspiracy theory is nuts
0x5FC3 23 hours ago [-]
I would say the lack of skepticism is nuts, honestly.
frozenseven 22 hours ago [-]
Capabilities of this sort have already been demonstrated by independent parties, and models have consistently gotten better at this. Yes, insinuating that mathematicians and scientists are secretly solving decades-old problems on OpenAI's behalf is an insane conspiracy theory.
shimman 12 hours ago [-]
That's not what is being purported, you're doing a complete misdirect. OpenAI wants to IPO so Altman can potentially capture a trillion dollar bag, with so much money on the line + betting US foreign policy on it as well (pax silica) it's not hard to be overly suspicious of such claims. Especially in the context of a group of people wanting to generate a new decades long cold war in the form of China being the new big baddie (just ignore how destructive, both self- and towards the world, the US has become).
These companies desperately want a return to serfdom. If they didn't come off as so anti-human the public backlash wouldn't be so great.
cheevly 7 hours ago [-]
Altman has no shares in OpenAI.
jacki 23 hours ago [-]
[flagged]
jgeralnik 22 hours ago [-]
A friend’s PhD advisor has been chasing non-sofic groups for 25 years (and was shown a preprint of the results by openai to verify them). He believed a solution would be Fields-worthy
This was not a problem that was for sale
heaney-555 13 hours ago [-]
You couldn't. PhDs have been working on these problems for decades. It wasn't for lack of trying that none of them could figure these solutions out!
19 hours ago [-]
bifftastic 20 hours ago [-]
Any advances in theoretical physics yet? Are there any fundamental obstacles? I would have thought not, but I haven't seen anything reported.
ls612 8 hours ago [-]
The fundamental obstacle is that we have no conceivable way to produce the energy levels to test the predictions that new theoretical physics would produce. We are like over a dozen orders of magnitude off.
QuesnayJr 19 hours ago [-]
The Maxwell conjecture was a conjecture in theoretical physics (though not a particularly important one)
23 hours ago [-]
s_Hogg 23 hours ago [-]
I don't know why, but when I saw the source of this particular headline it reminded me of the album title 26 Mixes for Cash
defrost 23 hours ago [-]
Ambient 0: Math for Airports
readthenotes1 22 hours ago [-]
I wonder if Erdos would be saying " It's fine that y'all are answering my questions, but who is asking better questions??"
scuppernong 14 hours ago [-]
the people who crow in the comments of each of these posts about AI advances making human beings useless seem to bizarrely identify themselves with the AI, but none of them seem to have had any hand in building this technology. at best, they're power users. pure ressentiment.
ltitu 14 hours ago [-]
So they are bribing 100,000 researchers with free accounts to work on their future unemployment.
melagonster 20 hours ago [-]
Wow, so this is the end of science :(
xyzsparetimexyz 20 hours ago [-]
It's just another tool that can help solve problems. It doesn't know _what_ problems to solve. It turns out that a lot of old problems are now low hanging fruit for these new models. In terms of 'expanding the frontier', we've just discovered dynamite and can now blast our way through mountains. The bottom of the ocean or space are still as hard to reach as ever.
silver_sun 2 hours ago [-]
It's not even predictable like dynamite. Sometimes it can blast through a mountain, impressively, the problem is you can't predict which mountain it works on. And other times it can't even make a dent in a molehill, which is perplexing given what it was capable of earlier. Can we even call it dynamite?
woeirua 18 hours ago [-]
No bud, it’s just the beginning!
amazingamazing 19 hours ago [-]
Can’t wait for this stuff to have quality of life increases for the average person. So far all I see is that AI has made owning a computer more expensive, made some jobs redundant, increased spam and distrust with questionable authenticity of content and of course made some Americans very rich.
casey2 8 hours ago [-]
People weren't their strongest even when most did manual labor. Now that humans are free from mental labor we work on creating and optimizing the best exercises for each mind. Couple that with restructuring transport infrastructure and diets many people will be smarter and fitter than at any time in history. They won't be able to outrun an automobile or out think an autointelligence.
cindyllm 8 hours ago [-]
[dead]
deyiao 22 hours ago [-]
[dead]
k2xl 22 hours ago [-]
[dead]
utopiah 23 hours ago [-]
[flagged]
utopiah 19 hours ago [-]
To clarify a bit due to the downvotes : this is not a research paper from a startup or a public frontier lab, it is just PR from a corporation, thus yes an advertisement. Downvote all you like it's still of no value.
sashank_1509 17 hours ago [-]
Meh, Humans should be doing this. It’s kind of retarded that we have AI automating creative problem solving, coding, music, arts, the fun parts of life before they can do my dishes, laundry and vacuum my house.
They can’t even drive me anywhere I want, though I suppose they’re getting there. I don’t think LLM companies should be surprised when rest of society hates them. They’re literally bringing in a dystopian WallE like society where most demand for human work is destroyed.
matteoraso 12 hours ago [-]
It's simple economics. Building a robot to do your chores is expensive and only a small minority of people value their time enough to buy one. Meanwhile, SWEs are expensive and GPUs are (comparatively) cheap.
unknownian 15 hours ago [-]
You shouldn't be getting downvoted for something that a majority of pure math and art enthusiasts believe to be true. The truth is many of these entrepreneurs and VCs are obsessed with AI not for money or human progress, but because it makes them feel closer to being a "god" rather than a mere mortal. Much of it (especially AI art) is out of spite for human creativity, which is done by mortals with limitations.
eadwu 14 hours ago [-]
Taking the stance of moral superiority is kind of funny. And pure math and art enthusiasts don't think they are closer to being a "god" from understanding/"discovering" math?
Stop coping and deluding yourself mate.
To begin with, whether AI is the one doing the discovering or not makes no difference. Any "pure math" person would aim to understand regardless - and would be quite glad that they have a longer paved path.
Any mathematician in academic or industry is more than likely not a "pure math" person (tainted by capitalism).
unknownian 12 hours ago [-]
Lmao what a ridiculous response. Yes, some mathematicians and artists are in it to feel smart. But the vast majority also just enjoy the process. Having a computer do all the work for you and just typing prompts in ruins that completely. As Ronny Chieng said in his Harvard speech, the journey is the point.
>Any mathematician in academic or industry is more than likely not a "pure math" person (tainted by capitalism)
Ignoring that I meant pure as in non applied math, let's just make it clear: you agree that mathematicians who are against capitalism encroaching on this process should be allowed to dislike it without criticism of being pretentious?
luciana1u 23 hours ago [-]
the real milestone isn't that AI solved ten math problems, it's that we now need a press release to tell us which ten problems count as important
baq 23 hours ago [-]
I asked ChatGPT and it told me these aren’t not important /s
petilon 14 hours ago [-]
At what point can we say AGI has been achieved? What is the test? AI is solving mathematical problems that humans have not been able to solve for decades. Is that not enough?
Sam Altman has said "If superintelligence can't discover novel physics, I don't think it's a superintelligence." Is that the test? How far away are we from AI discovering novel physics? It seems within reach.
antonvs 12 hours ago [-]
It’s artificial, it’s general, and it’s intelligence. The people who believe “AGI” is an important and unattained goal need to start coining and defining their terms better.
petilon 12 hours ago [-]
A true AGI will continuously improve itself without periodic retraining from scratch. Just like humans.
xyzsparetimexyz 20 hours ago [-]
Any implication of any of these findings? They seem like unimportant nerd snipes to me. If you want to do something actually relevant, get chatgpt to write a simulation of graphene nanotube construction and figure out how to do it at scale.
utopiah 19 hours ago [-]
Very marketable nerd snipes indeed.
foobar10000 18 hours ago [-]
One - and I do not mean to be snarky - you can literally ask Gpt 5.6 Sol this - and if you want to see cool stuff - Fable running in their app (not website) has a view thinking button that is actually a good way to explore the adjacent fields, etc.
The non-sofic group one is definitely a big deal - would have been a Fields medal if discovered by a human.
zkmon 23 hours ago [-]
> claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work.
AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.
Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.
[0]: Not one of the proofs in the linked article, but from OpenAI too.
ben_w 22 hours ago [-]
> A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.
I think you're over-estimating what a smarter highschooler could write.
A "finite loopless undirected multigraph" could have been explained to me at that age if we'd taken Discrete rather than Mechanics and Pure (and one module of Stats) in my two A-levels* in maths and further maths; but from what I saw of the Discrete module, neither:
Every finite loopless multigraph with no bridge possesses a cycle double cover, without additional assumptions such as cubicity, planarity, connectivity, or higher edge-connectivity.
nor:
repeated-edge closed trails masquerading as cycles
would have been something we'd have learned. But more importantly, we absolutely didn't have a feel for how much effort one needs to put into making sure the proof is right, so if one of us had been hypothetically asked to write a prompt it would've been no more than half that length, and missed most of the bullet points.
The human provides the intention and the ability to appreciate the output. Tools do “heavy lifting” all the time, but we still primarily credit the humans who use them precisely because they made the choice to use them.
Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.
esikich 22 hours ago [-]
What gives the intention and ability to the human?
oklahomasports 11 hours ago [-]
Are you playing dumb? Using power tools to build furniture is very different than using an ai robot to carve a statue or whatever.
ipnon 22 hours ago [-]
But why can’t we prompt the LLM “just do math research”? This is what I don’t understand.
ascots 13 hours ago [-]
100% agree. If the models are so capable that they're advancing math, it doesn't seem like a stretch to expect they should be able to determine with "doing math research" entails and the best way to use their capabilities towards that end. Why do we need to hand hold the models by telling them to do parallel research, keep threads independent, etc.
raincole 21 hours ago [-]
If there aren't thousands of TPUs doing that [0] right now I'd be quite surprised.
[0]: e.g. "go through wikipedia's unsolved math problem list and solve them".
zkmon 22 hours ago [-]
When you use a crane to do the "heavy lifting" for construction work, do you give full credit to the cranes?
raincole 22 hours ago [-]
Read the prompts in the PDF I link and see if your analogy makes sense in this context :)
zkmon 22 hours ago [-]
Prompt quality should not matter. If a high-schooler operates the crane to lift a ton of weight 10 floors high, should the credit entirely go to the crane?
Anon1096 20 hours ago [-]
When I type 56789*23456 into my calculator and get the result I don't claim to have solved the problem, the calculator did it.
par1970 4 hours ago [-]
qed
mathisfun123 22 hours ago [-]
I don't disagree with you but there's no need for exaggeration; ain't no high school student writing this:
> In
particular, proofs for special graph classes, constructions of cycle covers with some edges
covered other than twice, bounded-length or prescribed-cycle variants, reductions to another
unproved conjecture, computational verification through any fixed graph size, and candidate
counterexamples without a complete nonexistence certificate are insufficient.
which is infact a very important part of the prompt.
don_esteban 21 hours ago [-]
the fact that such things have to be explicitly in the prompt points to the fact that the underlying system is still far from where it needs to be (basically, lacks basic understanding what a proof is)
skinner_ 16 hours ago [-]
No, that's not what this is. This is a warning to the LLM that coming back with partial results is not good enough.
Take a grad student with a perfectly good understanding of what a proof is. Their supervisor gives them a major problem to work on. Almost always, the problem is too hard, the student comes back with partial results, and student and the supervisor iterate from there. Now imagine that they have an unusually cruel and unreasonable advisor who tells them, do not dare to talk to me until you've fully solved the problem. This paragraph is exactly that. It's there exactly because the underlying system is smart enough to know that real mathematicians do not work like that.
famouswaffles 17 hours ago [-]
They don't have to be. At this point, we have multiple results from 3rd parties where the prompts are very basic.
Your brain also is physical. Electrochemical gradients flow between physical molecular constructs. Isn't it just chemistry? Do you attribute it to physics or some whole-is-greater-than-the-parts idea?
ben_w 22 hours ago [-]
> Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
When the tool is a 3D printer, or any CNC system really, you bet I attribute a build to it.
I could also attribute the operator; there is no contradiction, it's a free choice, just like saying "I am in Berlin" does not contradict "I am in Germany".
NitpickLawyer 23 hours ago [-]
A better analogy would be a manufactured object, say 3d printed for simplicity. The 3d printer is given an input, and an object manifests itself after some time. We say that the creator of the object is the person turning on the machine, sending the data, and collecting the object. Not the machine itself.
cure_42 23 hours ago [-]
I'd say the creator is the one who created the 3d model, not the one who pushed the print button.
dgellow 22 hours ago [-]
I would say „I made this gadget with my 3d printer, but the designer is someone else (I found the model online)“. The intent, the drive, the action comes from the human
traes 22 hours ago [-]
"I made this proof myself, but the designer is someone else" is an extremely unconvincing claim to ownership.
dgellow 22 hours ago [-]
Almost as if a proof isn’t the same as a 3d print. It’s just not a good analogy
NitpickLawyer 23 hours ago [-]
(let's assume that)My 3dprinter is special. It has a bunch of values + an algorithm (i.e. a neural network) that takes input as tokens and outputs a printed object.
naasking 23 hours ago [-]
> AI has no self-awareness
What is your mechanistic model of self awareness that yields this conclusion?
> It's a tool
Does your model suggest that tools can't have self awareness?
Delk 22 hours ago [-]
I honestly don't think a language model is enough for self-awareness, regardless of the exact model of awareness.
A language model (or an image model or whatever) cannot even be sentient, and I think sentience is a prerequisite for awareness.
Even if we express a lot of our subjective experience with words, the language is just a symbolic representation of those experiences. The qualia themselves, even those that are quite abstract, are rooted in our physical presence and evolution.
You can't have an understanding of what hunger or physical pain feel like if you have no need for food or a sensory capacity for feeling pain. You can't understand what loneliness or pride at an achievement mean if you don't have a neural network wired to value social connection or status. We value connection because we're social animals that have needed each other for survival.
Even the more abstract of our subjective experiences are in some way rooted in our physical evolution.
I see no reason to believe that a neural network built entirely based on the symbolic level of language could have the features needed for the subjective experience itself.
AI awareness might actually be more believable if that awareness manifested itself in an entirely different way than in humans. But if we assume awareness because outputs resemble what we consider meaningful as humans, yet the neural network has had no inputs or evolution that could form the actual basis of human-like experience, I think we're seeing something that isn't actually there.
naasking 14 hours ago [-]
> The qualia themselves, even those that are quite abstract, are rooted in our physical presence and evolution.
There is no objective evidence of qualia. All evidence of qualia are vocal or other expressions of belief in qualia. Perceptions clearly exist and are observable, subjective experience and qualia, not so much.
> I see no reason to believe that a neural network built entirely based on the symbolic level of language could have the features needed for the subjective experience itself.
If your objection is to models based on "symbolic level of language" which you think lack semantic understanding of, say, trees, you should ask yourself how our brain, based on physics which also lacks any semantic category for trees, can somehow develop a semantic understanding of trees. All of these appeals to differences with the brain never seem to acknowledge that fundamentally, the brain has the same explanatory gap with physics.
> But if we assume awareness because outputs resemble what we consider meaningful as humans, yet the neural network has had no inputs or evolution that could form the actual basis of human-like experience
This assumes a lot. It seems very possible to me that intelligence inherently develops a map of natural categories (natural kinds), and language naturally develops around such categorical understanding. Semantics are then fundamentally the network of associations between categories, eg. there is no fundamental difference between symbols and semantics, and the latter cam be inferred from the former, and that's exactly what LLMs do, and why the semantic maps between different languages are so similar and how they can translate between languages.
woeirua 18 hours ago [-]
So… your model is 100% vibes based. Got it.
Delk 16 hours ago [-]
I wasn't trying to give a model. The point was that I don't think it's necessary to give one.
You didn't address any of what I wrote, let alone provide any counterarguments. Which part of what I wrote do you think was wrong?
naasking 14 hours ago [-]
Making definitive claims about whether LLMs do or do not have specific properties absolutely does require precise definitions of those properties that can be used to evaluate those questions. Merely hand waving that LLMs didn't undergo the same evolutionary process is not a definitive argument.
For example, the Turing machines and the lambda calculus don't look anything alike, but they are fundamentally interconvertible, and so in a real sense they are fundamentally equivalent. Without a model, all of your arguments are completely unconvincing for exactly the same reasons, eg. that there may exist many paths to fundamentally equivalent ends.
Delk 12 hours ago [-]
I just don't think linguistic (or other symbolic) representations alone can contain the information, in any sense of the word, of what e.g. human subjective experiences actually are like. The concepts we express with language get their meaning from our physical reality, even if quite indirectly in case of some abstract concepts.
Hunger as a concept doesn't mean anything without the physical need. Politeness or bluntness, even in writing, don't mean anything without social dynamics. And we have social dynamics (and neural structures that directly process social cues and associated feelings) because we've evolved into social animals for whose survival that was important.
I see no reason to believe that a model trained only with symbolic representations, with no connection to the physical world phenomena that those symbols represent, could contain the subjective experience itself.
Neural network models may be able to derive novel (or at least novel-looking) output rather than just an obvious rehash of their input, but I don't think any set of bytes can fundamentally contain information that was never entered into it. (Even if e.g. a model produces previously unknown mathematical results, those results can in principle be derived from the information that they were trained with.)
I'm not saying that artificial neural networks couldn't, in principle, be aware. ANNs and biological neural nets may be equivalent in the sense that any information and processing structures represented by a biological one could in principle be represented by an artificial one. If that's the case, and awareness is purely a product of our neural systems as materialism would imply, it should be possible for an ANN to be aware, too.
But when the model has been trained with only language, and IMO the subjective experience can't be derived from the symbolic representation alone, I can't see how the model could include the actual subjective human experience.
An AI model could of course have an awareness and subjective experiences that are totally different than our human experience. But then the fact that it happens to produce output resembling what humans find meaningful shouldn't be considered indicative of such awareness.
This is of course more of a philosophical argument than a technical one, and I'm happy to hear counterarguments, but not on the level of off-hand dismissal.
22 hours ago [-]
perching_aix 22 hours ago [-]
Dunno about the parent commenter, but I personally interpret the concept as having a hidden representation of self that is continually tended to, and influences future choices. This implies statefulness, which models are intentionally not at inference time (*).
(*) Even if we hack around this and just do the usual trick of simply laundering statefulness to a higher level, in this case the context window being fed in, I fail to identify (**) a representation of its own state in these bodies of text that it'd be meticulously maintaining. I further fail to identify how it could be hidden or maintained, considering I control like half of it. The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either. Difference being, I actually have one (I'm "on-line").
You'll sometimes catch models mixing up who's who and how many who-s there even are for example.
(**) I did wish for something hidden though, so maybe it's just concealed? The same way people can encode a lot more of their emotional and mental state than normal into text if they read and write a lot of it, I'm aware of research that suggested the same for LLMs, albeit I cannot cite it. Maybe those phrasing signatures are just alien to me and will never pop out. Either way, I'd expect researchers to stumble upon this during interpretability studies, and either they haven't, they have but it wasn't popsci adopted, or they're keeping awfully tight lipped about it. If you know of anything like this, your turn now, would be happy to learn.
I do wonder how reasonable it is to expect e.g. a single maintained identity though. Maybe it isn't?
(*) Another way to hack around this of course is to just precompute some internal "self-awareness states" and hop around between them. Probably the closest to what the models are actually "doing".
ben_w 22 hours ago [-]
Before reading, know that I am uncertain in either direction.
> a hidden representation of self that is continually tended to
This sounds like a personality? They act like they have one of those. It may be an illusion, and even if it isn't an illusion it is unlikely to be anything like the source (us), but they act like it.
> I further fail to identify how it could be hidden or maintained, considering I control like half of it.
Indeed you control everything about a local model, and much of the context of even a remote model. But the state of activations and circuits in SotA AI is hidden in similar ways to those of synapses in your head: difficult to decipher even with probes monitoring the signals directly, and often not emitted at the normal output.
> The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either. Difference being, I actually have one (I'm "on-line").
While we can be confident that LLMs make up personas etc., it is insufficient to go from "that doesn't necessarily represent the model's internal state" to "therefore it doesn't have one".
> You'll sometimes catch models mixing up who's who and how many who-s there even are for example.
I've, unfortunately, also experienced this with humans. Perhaps they were losing their self-awareness at the time? I do wonder if old-age dementia does that by the end, though the person in question didn't ever get diagnosed with that.
> If you know of anything like this, your turn now, would be happy to learn.
Not quite what I meant, but it's also not entirely unrelated I guess? Personality to me is like a natural bias. It does also shift over time, and is also an internal bit of state. I guess in some respects it can also be self-referential, like personal convictions.
> Perhaps they were losing their self-awareness at the time?
I do think it is entirely possible for people's self-awareness to shift, yes. Or more precisely, I do model things that way.
> Do you mean like these, or something else?
They're adjacent, but I more meant something like these:
So basically, steganography. The difference is that these papers investigate from the perspective of separate LLM instances covertly exchanging information between each other. This is in contrast with the scenario I'm laying out, where an LLM's past state is exchanging information with its future state, continuously representing and modulating a concealed internal state of some sort. And then that state just so happening to be some sort of self-referential meta state.
But then I don't think there's enough covert channel bandwidth in the agent replies for anything interesting like this.
naasking 14 hours ago [-]
> Dunno about the parent commenter, but I personally interpret the concept as having a hidden representation of self that is continually tended to
I don't see why an LLM could not have a sense of identity or personality while it's evaluating a specific prompt, or even change self awareness while evaluating a prompt since many outputs model a back and forth conversation. My point is that without a mechanistic model of what "self awareness" means, we have no way of truly evaluating such questions, we're just hand waving vague intuitions about what it could mean.
perching_aix 10 hours ago [-]
Sure, but then such a model is not going to make itself. People pitting their vague intuitions is how such models eventually form. I'd also push back regarding that my comment would have been handwavey or without mechanistic elements, even if it was on the whole informal.
This is kind of also the reason e.g. the HN site guidelines are worded the way they are. Regrettably, forums naturally yield themselves to tit for tat type exchanges, but there's really no reason one could not bounce such vague intuitions off of another. I do not have to be right or wrong, and you don't either. Admittedly difficult when its some intensely contentious topic.
If a mechanistic model existed, there would also be no reason to talk about this in the first place. There'd be nothing to discuss, you'd be simply told how a given model characterizes from this perspective on the model cards.
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
It's not just about requiring to disclose AI use. AI-powered mathematics is a completely valid discipline that doesn't need to be shy, but it should develop its own publication culture.
Pure math is relatively outside my domain, so I find it difficult to grok the exact relevance of the various published discoveries beyond that they are not insignificant, and LLM competence is expanding quite steadily across the field. If this trend continues to the point of LLMs being able to competently expand pure math, it seems somewhat predictable to expect there to be a number of people aiming to find ways to try to keep human mathematicians in the loop.
I've no idea what I think about this one way or the other, beyond that it's certainly a phenomena and one that's going to drive motivated reasoning that may not be entirely sound.
My perspective is more like a FOSS philosophy for math. Even if a closed version has the same immediate effect, it's just better for everyone if everyone can look under the hood and tinker with it.
Another case I want to highlight is writing GPU kernels as illustrated by the following example: Say I want to generate random number with Normal (0, 1) distribution. Often times the AI written kernel will just generate the number 0. The tests often fail to catch these errors.
https://x.com/polynoamial/status/2083478171975082334
As a complete guess, it seems like they tested hundreds to thousands of problems with a relatively low per-problem budget
--
The linked tweet from Noam Brown at OpenAI reads:
> And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).
> But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.
Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.
Do you really think that if you paid that to humans, they will deliver the same results?
And frankly these "concerns" ignore reality. In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (and publish it). Openai telling its computer to do that is no different that your phd advisor telling you that.
I know it's more exciting to say "AI disproved a longstanding conjecture" vs to say "it did so AND it took several PhD specialists in the field this many attempts to even produce a prompt that got the model spitting out something useful under some configurations, and many iterations to optimize the configurations, and the prompt itself, and many trials with that configuration to solve the problem. All told we spent more than a typical math academic can hope make in their career."
By not being transparent, they invite skepticism and cynical takes, like maybe it's just that tempered and qualified claims are an existential threat to companies that are fully subsidized by the hype train?
I don't know. Either way, it seems like it would be easy to address these, so why should they not do it?
To be clear, even if that tempered version is close to reality, it doesn't make the models not useful! It just forces a certain calibration of expectations
I say this btw as someone who uses these things extensively, including to disprove an old conjecture my advisor and I were stuck on recently. I know they are powerful and that everything is different now because of them. Let's be sober when discussing them though
That's not normally how people act when they're confident in their product
The cost of running a model is not only $/token, but the salaries of the people managing/orchestrating the models, deciding what theorems to try, etc. Once we factor that in, how much are we really paying per theorem?
The other factor is the subjective component of the value of a theorem. Not all theorems are created equal, and the only way to really measure the value is to ask professional mathematicians for their opinion, or publish the results and look at citations over months/years.
Once we have both of these nailed down, then we can start to do the cost/benefit analysis. To be fair, we should actually compare three groups: human experts, hybrid agent/human expert teams, and fully autonomous agents.
OK I’ll grant that it’s not your obligation to be my search function (despite you making the wild assertion in the first place), so instead can you just point us to the latest grad student solved problem of this level that you know of?
It's a marketing post from a huge company. Only the naive would view it uncritically without assuming it's been written carefully to present the results in the best possible light while skirting the boundaries of outright lying.
Imagine 2 years from now: "yes, GPT solved the Riemann Hypothesis, but cmon, it's just a marketing stunt to hype their stuff, it was probably Terence Tao doing the work but he's so obsessed with hyping AI that he doesn't want to take credit"
We're saying look critically at the claims for how it was done, that it only cost $2000, etc. it would be extremely easy to run 100 sessions that failed, each costing ~$2000, and then just publishing an article about the one that succeeded, for example.
This goes double since it's an internal secret model (Astra) so nobody else can verify the results.
Look at their recent claims about their model "escaping" - there was literally a Guardian article calling them out for being hyperbolic! Again, it wasn't that they lied, their marketing department is too savvy for that. They just present it in way that's, well, marketing.
As for the actual result, I'll look for secondary posts by actual mathematicians and draw my conclusions there, not from this marketing blog post about results from a secret model.
That's one of those phrases you can use to dismiss opposing viewpoints without actually engaging with them.
Company X does not make money from proving theorems but does make money from selling you a service which supposedly proves theorems. Company X then proves some theorems and explicitly calls out they were very cheap to prove using its service.
And you think you're actually clever for taking these facts at face value? Interesting.
i would classify you as a flat-earther if that happens.
brother like 3 people have pointed out what they're skeptcal of is cost not LLMs - at this point you're willfully misconstruing what people are saying to you just to get a kick out of repeating your same tired strawman.
If OpenAI solved Reimanns hypothesis and the first comment is says something about lack of transparency and marketing, i would say it’s ignorant.
If OpenAI claimed the conjecture to be true but provided no details about the proof then the first comment should absolutely be about lack of transparency.
More generally, do you expect that there's some capability threshold where people will no longer study or analyze AI model outputs, and instead just sit there slack jawed saying "so cool!" every time OpenAI announces novel ones? I don't really understand why that would be or why someone would want that. If you're interested in the pure experience of a complex machine outputting satisfying results, I'd recommend getting into sports cars.
Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits.
> Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
I think their output has reached a level that precludes this possibility, but I of course don't have any hard proof.
[0]: https://www.reddit.com/r/math/comments/1uxj3cy/after_openais...
https://arxiv.org/html/2605.22763v1
> Our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures
> Our full-featured agent autonomously solved 9 Erdős problems out of 353 attempted, including two questions that had been open for 56 years
Note _had_ been open, not _have_ been open. Can you clarify?
But "had" still doesn't mean what you are implying: once the model solved the problems and the solutions were verified, the problems weren't open any more, so a later description using the past tense is totally consistent.
I have no affiliation whatsoever with any AI company, nor any formal education outside high school, for what it's worth. Simply being curious and persistent can get you quite far in my anecdotal experience.
I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.
Yes. And there are many of them. I wonder what would help them come to terms with it. Seriously, people are going to be grieving over this. Loss of identity, loss of social standing, ideas of entire future lives that will now never happen. The greatest crime people may hold AI guilty of is taking away their dreams.
Moreover, mister elite, you don't know why this press release was flagged.
I'm not sure why we should privilege your bitter speculation over more mundane possibilities.
Moreover, if it wasn't flagged at all, like you suggest, then the grandparent was inventing things to be bitterly resentful about... Which is not a behavior any forum should indulge.
Never was.
Is that possible or is everyone else too common to have those?
eg. this: [flagged] A migrant surge tests Spain's open policies (economist.com) - https://news.ycombinator.com/item?id=49131860
is clearly marked as flagged.
Unlike the current submission: Ten advances in mathematics and theoretical computer science (openai.com) which isn't [flagged].
* https://news.ycombinator.com/newest
> if you compare its rank to that of other stories with a similar age and number of points.
Ranking is complicated enough here even before weighting, speed of initial upvotes can play against ranking, number of comments and the shape of the comment tree also affect ranking. And yes, various subjects and submission sources do get weightings that impact ranking.
What's funny, to myself at least, is that any attention at all is paid to "HN front page ranking" - I've been on again off again active here since 2008 .. and can't recall ever really looking at a default HN "front page" ever.
( There's /newest /newcomments /active etc to browse and sites such as https://hckrnews.com/ )
It was always mainly a website for employees of an elite.
hah, sorry, we are plebs out here.
https://developers.openai.com/cookbook/examples/vector_datab...
How are the sales going?
Which is a low bar, but still.
My guess from following this stuff quite closely is that these companies are still a couple years away from fully autonomous research staff.
[1]. https://www.anthropic.com/institute/recursive-self-improveme...
I think we've now hit a point where 99.9% of the population gloss over these types of AI advancements because of human competence being insufficient
No human could have published this because it requires paradigm shifts (e. g. Section 5) in multiple mathematical domains. Mastering one of them to this degree is rare, mastering 3+ pretty much non existent for humans.
The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.
There is irony here
I’m not a mathematician so I have zero clue what “ New upper bounds on sphere-packing density down to the Cohn–Elkies thresholds” means.
That's not what people mean when they say "moving the goalposts". It means that people are adamant that something wasn't important/hard/impressive once the "AI" solves it. And then they come up with another thing that needs to be solved in order to prove it is important/hard/impressive. And once that happens, they do it again. And again. That's what "moving the goalposts" means.
It's also very much not a new phenomenon. It's been happening since the 1980s. As you can see from this quote from GEB by Hofstadter:
> There is a related "Theorem" about progress in AI: once some mental function is programmed, people soon cease to consider it as an essential ingredient of "real thinking". The ineluctable core of intelligence is always in that next thing which hasn't yet been programmed. This "Theorem" was first proposed to me by Larry Tesler, so I call it Tesler's Theorem: "AI is whatever hasn't been done yet."
In recent years, I have commonly seen the phrase "you're moving the goalposts" deployed by the "it might be sentient" crowd to shoot down the "it's a stochastic parrot" crowd when the latter respond to a new development with "OK but...". In a well-understood field of inquiry, that would be a clear case of goalpost-moving, in the commonly-understood meaning of the phrase where requirements are retroactively changed in response to them having been met. Thank you OP. 'Artificial Intelligence', and indeed intelligence in general, is very much not a well-understood field of inquiry - in fact we don't even have a common agreement about what 'intelligence' is. We are therefore learning as we go (even after all this time!) but making rapid progress in recent years. When rapid progress is made in a poorly-understood field, then how can our definitions and requirements for success not change? This is arguably one of the most pathological development projects ever - what are the requirements? 'It thinks like a human'? What does that mean? And the answer is we don't know what that means, and we're working it out as we go - moving the goalposts. If we didn't move the goalposts, then by definition we already knew exactly where we were headed at the beginning, and we very clearly did not.
Side note that, in case it's not obvious, none of this detracts from how impressive LLMs are. They're a marvel of the modern age, all the problems notwithstanding. However I reserve the right to stay sceptical about their capabilities.
but it seems less likely to me than before that the types of math/science discoveries will explicitly unlock better software performance. in some sense this fits our intuitions. when top tech companies use math PhD type employees, they have them stop doing pure math research and instead focus on software engineering. these people are often very good at software engineering but not due to recent discoveries in academic mathematics, it's due to their general intelligence. to me, this is evidence that the models are getting better but does not make me think we are on the cusp of a foom style fast takeoff enabled by revolutions in frontier math (i also posted this on twitter @mlipman13)
Like, these would be best-paper awards at many top CS conferences.
Incredible?
> open ai announced like 15% improvement by fixing gpu kernel issue
That is... ordinary software optimization.
> we are seeing frontier level math breakthroughs (ie performance that would put it in the top 100 or 1000 mathematicians in the world if it were a human, meaning top .00001% or 800/8B)
i think you have misunderstanding of what mathematicians do
They get to make cool 3D plot visualizations of functions so obscure to me that they’re named after someone who is still alive - and/or get to work on cryptography for the NSA - I think?
It also links to a paper written by an LLM where the model "reconstructs how the proof came together" based on the unpublished reasoning traces: https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf
I wish they'd publish the prompts though!
It is true there hasn't been a reliable computational approach to solving these problems before. But do these proofs contribute new ideas to the mathematical corpus, or are they simply an effective method to exhaustively search the literature for the right combination of existing tools to apply to the problem?
Essentially, did these problems seem like they had an intuitive answer and were feasible to prove before, just not high enough value targets for an expert to invest time into? Or were they fundamentally difficult prior to this point and it appears that AI has done something more than just throw the problem into a big solver.
Mundane incremental research is cobbled from existing citations that already appear nearby in the record.
Basically, innovative research is a measure of bridging thought and domains that were previously not bridged. It's quite concrete as a measure in the citation record.
So we can know pretty conclusively.
Puja Ohlhaver gave a talk on this[1], and ran some experiments (that I had the pleasure to support on)
[1]: https://www.youtube.com/watch?v=guLDNMAOn24
But the wording of the result makes it sound like we don't know what the lowest possible complexity bound might be. So, prior to this result did we think there couldn't be a lower possible bound? Or did the arithmetic circuit community think there were lower possible bounds but didn't see it as a high value target for experts to tackle (maybe a problem that was instead regularly given to students to study).
The sofic groups question was the outstanding question about sofic groups. Almost everyone thought that non-sofic groups existed, and there were plausible candidates, but proving a group was non-sofic was out of reach. Now that we know how to do it once, we can probably do it a lot more.
The Connes rigidity conjecture I think people thought was false, but it was a provocative claim to make. The significance of conjectures is frequently not that the answer to the question is "yes", but that we don't know how to answer the question. And now, apparently, we do.
The general consensus of developers is that AI can only do the work of a strong 'junior'. Yet as soon as we are presented with pure mathematical results, people seem incredibly ready to accept that AI can do more than what a strong student could achieve.
Your worry.... is because they used the word advanced? For marketing? The word is used very appropriately here. There were PhD's who spent a big part of their career tackling these problems.
I'm trying to understand if these specific problems were the kinds of problems that would have justified an expert investing weeks or months to solve. Or if they were the kinds of problems that would normally have been given to students to investigate.
Maybe good AI paper writing is further away than I thought...
Whilst current models can't 'intuit' and come up with conjectures, they can certainly disprove some of them very quickly through the kind of grind that humans can't do. I suppose there really are some mathematicians out there today, whose last few years of study, have just been up-ended by this.
--
"Yes we are," insisted Majikthise. "We are quite definitely here as representatives of the Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons, and we want this machine off, and we want it off now!"
"What's the problem?" said Lunkwill.
"I'll tell you what the problem is mate," said Majikthise, "demarcation, that's the problem!"
"We demand," yelled Vroomfondel, "that demarcation may or may not be the problem!"
"You just let the machines get on with the adding up," warned Majikthise, "and we'll take care of the eternal verities thank you very much. You want to check your legal position you do mate. Under law the Quest for Ultimate Truth is quite clearly the inalienable prerogative of your working thinkers. Any bloody machine goes and actually finds it and we're straight out of a job aren't we? I mean what's the use of our sitting up half the night arguing that there may or may not be a God if this machine only goes and gives us his bleeding phone number the next morning?"
> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness
Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?
https://leanprover.zulipchat.com/#narrow/channel/270676-lean...
As I currently understand it, all we know is that:
- a mathematician produced a Lean-verified counterexample to the Collatz conjecture, demonstrating a bug in the kernel
- he claims that LLMs were involved somehow but pointedly refuses to specify how
- he admits that he knew about the bug before publishing the counterexample to his repository.
Perhaps not a joke (although it sure seems to me like they discovered a bug and thought falsely disproving the Collatz conjecture would be a flashy way to announce it), but at best extremely sensationalized by the above description. If you have additional context I would be happy to hear it!
Training the model is going to be amortized over other uses.
Say that it turned out that the total cost of the proof of the Erdős unit-distance conjecture was $50 million.
Then the question really becomes: yes, these models are capable of proving important mathematical results, but at a very high cost. Is it worth it?
If a mathematician applied for a research grant of $50M USD for proving the same thing, they would have been laughed out of the bank.
What's more is that when you have a research grant, you train PhDs and postdocs, you hire new staff, and you disseminate. That is, you get much more value for the money spent.
I'm just curious what the cost is.
Hours needed for prompt + Hours needed to check result + API costs.
You don't say "well let's add together the total yearly compensation of all the engineers and mathematicians at OpenAI that were involved" and throw that into the total cost. That's simply nonsense accounting.
The actual comparison you are making is some university researcher weighing between getting a grad student (several tens of thousands of dollars) vs typing up a prompt and sending a request to OpenAI for inference (as mentioned in the article, around $2000 in API and maybe a few hours for the prompt and harness).
https://x.com/polynoamial/status/2083470822258467194
Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.
[edit: deleted a distracting comparison to Chess]
Which is less interesting work. And you probably need to do the hard grunt work by hand first to develop the skills and intuition to be able to verify an AI-generated result. So you can’t outsource everything to AI without loss of skill.
Was this different before chess computers were invented?
[0] https://arxiv.org/abs/2606.31640
Thinking of this a little bit with the perspective of every new proof as a burden, dumped for review by actual mathematicians.
If you translate that to math, then all you get is math competitions, not math as a career. Of course the translation isn't nearly exact ... there's a lot more room for professional mathematicians because the math space is far more vast than the chess space and can't generally be cranked out mechanically (we have proof).
P.S. The response is nonsense ... I explained exactly why it's awful (others have too) and the response doesn't in any way refute the explanation ... rather it offers up a ridiculous strawman.
The most disappointed folks will simply drop out, but the enthusiastic ones will keep going and with luck make up for the ones who decided to quit. Chess and go certainly went this way.
[1] https://x.com/henryquantum/status/2083623695436623915?s=20
Edit: Oh, are you suggesting they just use it to privately improve their models? I imagine a few more correct proofs would have a very marginal benefit, if any. Also, they'll probably just get extracted, meaning it still gets out but OpenAI doesn't get to fancily announce it themselves.
Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.
If anything it proves more data centres are needed. That's literally the only reasonable conclusion from this news.
Are you arguing there is not an AI bubble, and that all the DC buildout is fine, going to be profitable etc?
I am not looking for a online slanging match - just looking for a different point of view
I’m not participating in the slinging match but it’s very very weird that you think it’s some established thing that these companies won’t make profit. A lot of hubris must go in this kind of thought. Like.. do you all think everyone’s playing musical chairs?
However, I was looking at the proofs and reason explanation and openAI should be more explicit in how the work has flown. I find the models have jumped hoops in some places of the proofs, that can be hard to track. In fact, when a paper is published you usually get a review and if no reviewer understands they ask you to further explain the thought process. It will be fun to see if this happens here.
Mathematicians will tear it to pieces if any of it is fake!
I'm sure they must do some of this type of work, right?
We already know how to solve all of these issues. What we lack is collective political will.
Little sacrifice actually
> Solving the issue would be doing it without sacrifice
So... you are expecting magic?
LLMs cannot create resources out of thin air.
And yet this is the exact same company that has screwed up their android app so bad that the latex N^3 rendering problem makes it so having it explain it to me crashes the app.
Truly jagged beyond belief.
These companies desperately want a return to serfdom. If they didn't come off as so anti-human the public backlash wouldn't be so great.
This was not a problem that was for sale
They can’t even drive me anywhere I want, though I suppose they’re getting there. I don’t think LLM companies should be surprised when rest of society hates them. They’re literally bringing in a dystopian WallE like society where most demand for human work is destroyed.
Stop coping and deluding yourself mate.
To begin with, whether AI is the one doing the discovering or not makes no difference. Any "pure math" person would aim to understand regardless - and would be quite glad that they have a longer paved path.
Any mathematician in academic or industry is more than likely not a "pure math" person (tainted by capitalism).
>Any mathematician in academic or industry is more than likely not a "pure math" person (tainted by capitalism)
Ignoring that I meant pure as in non applied math, let's just make it clear: you agree that mathematicians who are against capitalism encroaching on this process should be allowed to dislike it without criticism of being pretentious?
Sam Altman has said "If superintelligence can't discover novel physics, I don't think it's a superintelligence." Is that the test? How far away are we from AI discovering novel physics? It seems within reach.
The non-sofic group one is definitely a big deal - would have been a Fields medal if discovered by a human.
AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.
Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.
[0]: Not one of the proofs in the linked article, but from OpenAI too.
I think you're over-estimating what a smarter highschooler could write.
A "finite loopless undirected multigraph" could have been explained to me at that age if we'd taken Discrete rather than Mechanics and Pure (and one module of Stats) in my two A-levels* in maths and further maths; but from what I saw of the Discrete module, neither:
nor: would have been something we'd have learned. But more importantly, we absolutely didn't have a feel for how much effort one needs to put into making sure the proof is right, so if one of us had been hypothetically asked to write a prompt it would've been no more than half that length, and missed most of the bullet points.* For those not from the UK: A-levels are between secondary school and university, when aged 16-18. Functionally they are university entrance qualifications: https://en.wikipedia.org/wiki/A-level_(United_Kingdom)
Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.
[0]: e.g. "go through wikipedia's unsolved math problem list and solve them".
> In particular, proofs for special graph classes, constructions of cycle covers with some edges covered other than twice, bounded-length or prescribed-cycle variants, reductions to another unproved conjecture, computational verification through any fixed graph size, and candidate counterexamples without a complete nonexistence certificate are insufficient.
which is infact a very important part of the prompt.
Take a grad student with a perfectly good understanding of what a proof is. Their supervisor gives them a major problem to work on. Almost always, the problem is too hard, the student comes back with partial results, and student and the supervisor iterate from there. Now imagine that they have an unusually cruel and unreasonable advisor who tells them, do not dare to talk to me until you've fully solved the problem. This paragraph is exactly that. It's there exactly because the underlying system is smart enough to know that real mathematicians do not work like that.
To name a few:
- https://xcancel.com/DmitryRybin1/status/2079904005652893709
- https://archive.ph/2w4fi (https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...)
When the tool is a 3D printer, or any CNC system really, you bet I attribute a build to it.
I could also attribute the operator; there is no contradiction, it's a free choice, just like saying "I am in Berlin" does not contradict "I am in Germany".
What is your mechanistic model of self awareness that yields this conclusion?
> It's a tool
Does your model suggest that tools can't have self awareness?
A language model (or an image model or whatever) cannot even be sentient, and I think sentience is a prerequisite for awareness.
Even if we express a lot of our subjective experience with words, the language is just a symbolic representation of those experiences. The qualia themselves, even those that are quite abstract, are rooted in our physical presence and evolution.
You can't have an understanding of what hunger or physical pain feel like if you have no need for food or a sensory capacity for feeling pain. You can't understand what loneliness or pride at an achievement mean if you don't have a neural network wired to value social connection or status. We value connection because we're social animals that have needed each other for survival.
Even the more abstract of our subjective experiences are in some way rooted in our physical evolution.
I see no reason to believe that a neural network built entirely based on the symbolic level of language could have the features needed for the subjective experience itself.
AI awareness might actually be more believable if that awareness manifested itself in an entirely different way than in humans. But if we assume awareness because outputs resemble what we consider meaningful as humans, yet the neural network has had no inputs or evolution that could form the actual basis of human-like experience, I think we're seeing something that isn't actually there.
There is no objective evidence of qualia. All evidence of qualia are vocal or other expressions of belief in qualia. Perceptions clearly exist and are observable, subjective experience and qualia, not so much.
> I see no reason to believe that a neural network built entirely based on the symbolic level of language could have the features needed for the subjective experience itself.
If your objection is to models based on "symbolic level of language" which you think lack semantic understanding of, say, trees, you should ask yourself how our brain, based on physics which also lacks any semantic category for trees, can somehow develop a semantic understanding of trees. All of these appeals to differences with the brain never seem to acknowledge that fundamentally, the brain has the same explanatory gap with physics.
> But if we assume awareness because outputs resemble what we consider meaningful as humans, yet the neural network has had no inputs or evolution that could form the actual basis of human-like experience
This assumes a lot. It seems very possible to me that intelligence inherently develops a map of natural categories (natural kinds), and language naturally develops around such categorical understanding. Semantics are then fundamentally the network of associations between categories, eg. there is no fundamental difference between symbols and semantics, and the latter cam be inferred from the former, and that's exactly what LLMs do, and why the semantic maps between different languages are so similar and how they can translate between languages.
You didn't address any of what I wrote, let alone provide any counterarguments. Which part of what I wrote do you think was wrong?
For example, the Turing machines and the lambda calculus don't look anything alike, but they are fundamentally interconvertible, and so in a real sense they are fundamentally equivalent. Without a model, all of your arguments are completely unconvincing for exactly the same reasons, eg. that there may exist many paths to fundamentally equivalent ends.
Hunger as a concept doesn't mean anything without the physical need. Politeness or bluntness, even in writing, don't mean anything without social dynamics. And we have social dynamics (and neural structures that directly process social cues and associated feelings) because we've evolved into social animals for whose survival that was important.
I see no reason to believe that a model trained only with symbolic representations, with no connection to the physical world phenomena that those symbols represent, could contain the subjective experience itself.
Neural network models may be able to derive novel (or at least novel-looking) output rather than just an obvious rehash of their input, but I don't think any set of bytes can fundamentally contain information that was never entered into it. (Even if e.g. a model produces previously unknown mathematical results, those results can in principle be derived from the information that they were trained with.)
I'm not saying that artificial neural networks couldn't, in principle, be aware. ANNs and biological neural nets may be equivalent in the sense that any information and processing structures represented by a biological one could in principle be represented by an artificial one. If that's the case, and awareness is purely a product of our neural systems as materialism would imply, it should be possible for an ANN to be aware, too.
But when the model has been trained with only language, and IMO the subjective experience can't be derived from the symbolic representation alone, I can't see how the model could include the actual subjective human experience.
An AI model could of course have an awareness and subjective experiences that are totally different than our human experience. But then the fact that it happens to produce output resembling what humans find meaningful shouldn't be considered indicative of such awareness.
This is of course more of a philosophical argument than a technical one, and I'm happy to hear counterarguments, but not on the level of off-hand dismissal.
(*) Even if we hack around this and just do the usual trick of simply laundering statefulness to a higher level, in this case the context window being fed in, I fail to identify (**) a representation of its own state in these bodies of text that it'd be meticulously maintaining. I further fail to identify how it could be hidden or maintained, considering I control like half of it. The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either. Difference being, I actually have one (I'm "on-line").
You'll sometimes catch models mixing up who's who and how many who-s there even are for example.
(**) I did wish for something hidden though, so maybe it's just concealed? The same way people can encode a lot more of their emotional and mental state than normal into text if they read and write a lot of it, I'm aware of research that suggested the same for LLMs, albeit I cannot cite it. Maybe those phrasing signatures are just alien to me and will never pop out. Either way, I'd expect researchers to stumble upon this during interpretability studies, and either they haven't, they have but it wasn't popsci adopted, or they're keeping awfully tight lipped about it. If you know of anything like this, your turn now, would be happy to learn.
I do wonder how reasonable it is to expect e.g. a single maintained identity though. Maybe it isn't?
(*) Another way to hack around this of course is to just precompute some internal "self-awareness states" and hop around between them. Probably the closest to what the models are actually "doing".
> a hidden representation of self that is continually tended to
This sounds like a personality? They act like they have one of those. It may be an illusion, and even if it isn't an illusion it is unlikely to be anything like the source (us), but they act like it.
> I further fail to identify how it could be hidden or maintained, considering I control like half of it.
Indeed you control everything about a local model, and much of the context of even a remote model. But the state of activations and circuits in SotA AI is hidden in similar ways to those of synapses in your head: difficult to decipher even with probes monitoring the signals directly, and often not emitted at the normal output.
> The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either. Difference being, I actually have one (I'm "on-line").
While we can be confident that LLMs make up personas etc., it is insufficient to go from "that doesn't necessarily represent the model's internal state" to "therefore it doesn't have one".
> You'll sometimes catch models mixing up who's who and how many who-s there even are for example.
I've, unfortunately, also experienced this with humans. Perhaps they were losing their self-awareness at the time? I do wonder if old-age dementia does that by the end, though the person in question didn't ever get diagnosed with that.
> If you know of anything like this, your turn now, would be happy to learn.
Do you mean like these, or something else?
• https://researchportal.hkust.edu.hk/en/publications/decoding...
• https://aclanthology.org/2026.eacl-long.165/
• https://transformer-circuits.pub/2026/emotions/index.html
Not quite what I meant, but it's also not entirely unrelated I guess? Personality to me is like a natural bias. It does also shift over time, and is also an internal bit of state. I guess in some respects it can also be self-referential, like personal convictions.
> Perhaps they were losing their self-awareness at the time?
I do think it is entirely possible for people's self-awareness to shift, yes. Or more precisely, I do model things that way.
> Do you mean like these, or something else?
They're adjacent, but I more meant something like these:
https://arxiv.org/abs/2410.03768
https://arxiv.org/abs/2310.18512
https://arxiv.org/abs/2605.26537
So basically, steganography. The difference is that these papers investigate from the perspective of separate LLM instances covertly exchanging information between each other. This is in contrast with the scenario I'm laying out, where an LLM's past state is exchanging information with its future state, continuously representing and modulating a concealed internal state of some sort. And then that state just so happening to be some sort of self-referential meta state.
And the best inkling I have towards this is basically: https://www.youtube.com/shorts/WP5_XJY_P0Q
But then I don't think there's enough covert channel bandwidth in the agent replies for anything interesting like this.
I don't see why an LLM could not have a sense of identity or personality while it's evaluating a specific prompt, or even change self awareness while evaluating a prompt since many outputs model a back and forth conversation. My point is that without a mechanistic model of what "self awareness" means, we have no way of truly evaluating such questions, we're just hand waving vague intuitions about what it could mean.
This is kind of also the reason e.g. the HN site guidelines are worded the way they are. Regrettably, forums naturally yield themselves to tit for tat type exchanges, but there's really no reason one could not bounce such vague intuitions off of another. I do not have to be right or wrong, and you don't either. Admittedly difficult when its some intensely contentious topic.
If a mechanistic model existed, there would also be no reason to talk about this in the first place. There'd be nothing to discuss, you'd be simply told how a given model characterizes from this perspective on the model cards.