Ten advances in mathematics and theoretical computer science
Posted by milkshakes 11 hours ago
Comments
Comment by aabhay 11 hours ago
I want to know:
1. How many total problems were given to the model, and what percent were left unsolved at what cost before giving up? 2. How many attempts did you give the model at solving these problems? 3. How expensive was the harness, e.g. did the model have access to a job cluster?
Comment by c7b 2 hours ago
It's not just about requiring to disclose AI use. AI-powered mathematics is a completely valid discipline that doesn't need to be shy, but it should develop its own publication culture.
Comment by jsenn 1 hour ago
Comment by c7b 1 hour ago
Comment by somenameforme 1 hour ago
Pure math is relatively outside my domain, so I find it difficult to grok the exact relevance of the various published discoveries beyond that they are not insignificant, and LLM competence is expanding quite steadily across the field. If this trend continues to the point of LLMs being able to competently expand pure math, it seems somewhat predictable to expect there to be a number of people aiming to find ways to try to keep human mathematicians in the loop.
I've no idea what I think about this one way or the other, beyond that it's certainly a phenomena and one that's going to drive motivated reasoning that may not be entirely sound.
Comment by jsenn 55 minutes ago
Comment by SpicyLemonZest 1 hour ago
Comment by black_knight 1 hour ago
Comment by lkirk 1 hour ago
Comment by c7b 26 minutes ago
Comment by wrsh07 58 minutes ago
https://x.com/polynoamial/status/2083478171975082334
As a complete guess, it seems like they tested hundreds to thousands of problems with a relatively low per-problem budget
--
The linked tweet from Noam Brown at OpenAI reads:
> And yes we did try other major problems without success. Sadly no Millennium Prize problems (yet).
> But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.
Comment by moscoe 7 minutes ago
Comment by whattheheckheck 5 hours ago
Comment by dist-epoch 8 hours ago
Even if the cost was $1 mil for these 10 problems, that's maybe 10-20 math researchers for a year.
Do you really think that if you paid that to humans, they will deliver the same results?
Comment by uh_uh 7 hours ago
Comment by dgacmu 4 hours ago
Comment by halJordan 1 hour ago
And frankly these "concerns" ignore reality. In any research phd course you're actively told to bite off something small and likely to be provable so that you can prove it (and publish it). Openai telling its computer to do that is no different that your phd advisor telling you that.
Comment by crazylogger 3 hours ago
Comment by vector_spaces 4 hours ago
I know it's more exciting to say "AI disproved a longstanding conjecture" vs to say "it did so AND it took several PhD specialists in the field this many attempts to even produce a prompt that got the model spitting out something useful under some configurations, and many iterations to optimize the configurations, and the prompt itself, and many trials with that configuration to solve the problem. All told we spent more than a typical math academic can hope make in their career."
By not being transparent, they invite skepticism and cynical takes, like maybe it's just that tempered and qualified claims are an existential threat to companies that are fully subsidized by the hype train?
I don't know. Either way, it seems like it would be easy to address these, so why should they not do it?
To be clear, even if that tempered version is close to reality, it doesn't make the models not useful! It just forces a certain calibration of expectations
I say this btw as someone who uses these things extensively, including to disprove an old conjecture my advisor and I were stuck on recently. I know they are powerful and that everything is different now because of them. Let's be sober when discussing them though
Comment by ifwinterco 3 hours ago
That's not normally how people act when they're confident in their product
Comment by fasterik 2 hours ago
The cost of running a model is not only $/token, but the salaries of the people managing/orchestrating the models, deciding what theorems to try, etc. Once we factor that in, how much are we really paying per theorem?
The other factor is the subjective component of the value of a theorem. Not all theorems are created equal, and the only way to really measure the value is to ask professional mathematicians for their opinion, or publish the results and look at citations over months/years.
Once we have both of these nailed down, then we can start to do the cost/benefit analysis. To be fair, we should actually compare three groups: human experts, hybrid agent/human expert teams, and fully autonomous agents.
Comment by wbl 3 hours ago
Comment by robotpepi 6 hours ago
Comment by kevinwang 6 hours ago
Comment by tchalla 5 hours ago
Comment by mungaihaha 7 hours ago
Comment by mirzap 7 hours ago
Comment by gbnwl 3 hours ago
OK I’ll grant that it’s not your obligation to be my search function (despite you making the wild assertion in the first place), so instead can you just point us to the latest grad student solved problem of this level that you know of?
Comment by whattheheckheck 5 hours ago
Comment by irthomasthomas 7 hours ago
Comment by simianwords 10 hours ago
Comment by traes 10 hours ago
Comment by simianwords 10 hours ago
Comment by esperent 10 hours ago
It's a marketing post from a huge company. Only the naive would view it uncritically without assuming it's been written carefully to present the results in the best possible light while skirting the boundaries of outright lying.
Comment by dist-epoch 8 hours ago
Imagine 2 years from now: "yes, GPT solved the Riemann Hypothesis, but cmon, it's just a marketing stunt to hype their stuff, it was probably Terence Tao doing the work but he's so obsessed with hyping AI that he doesn't want to take credit"
Comment by esperent 7 hours ago
We're saying look critically at the claims for how it was done, that it only cost $2000, etc. it would be extremely easy to run 100 sessions that failed, each costing ~$2000, and then just publishing an article about the one that succeeded, for example.
This goes double since it's an internal secret model (Astra) so nobody else can verify the results.
Comment by simianwords 6 hours ago
Comment by esperent 5 hours ago
Look at their recent claims about their model "escaping" - there was literally a Guardian article calling them out for being hyperbolic! Again, it wasn't that they lied, their marketing department is too savvy for that. They just present it in way that's, well, marketing.
As for the actual result, I'll look for secondary posts by actual mathematicians and draw my conclusions there, not from this marketing blog post about results from a secret model.
Comment by simianwords 2 hours ago
Comment by esperent 1 hour ago
That's one of those phrases you can use to dismiss opposing viewpoints without actually engaging with them.
Comment by SpicyLemonZest 2 hours ago
More generally, do you expect that there's some capability threshold where people will no longer study or analyze AI model outputs, and instead just sit there slack jawed saying "so cool!" every time OpenAI announces novel ones? I don't really understand why that would be or why someone would want that. If you're interested in the pure experience of a complex machine outputting satisfying results, I'd recommend getting into sports cars.
Comment by nxpnsv 9 hours ago
Comment by einpoklum 10 hours ago
Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
Comment by energy123 8 hours ago
Comment by brighteyes 4 hours ago
https://arxiv.org/html/2605.22763v1
> Our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars, proved 44/492 OEIS conjectures
Comment by einpoklum 4 hours ago
> Our full-featured agent autonomously solved 9 Erdős problems out of 353 attempted, including two questions that had been open for 56 years
Note _had_ been open, not _have_ been open. Can you clarify?
Comment by jsnell 4 hours ago
But "had" still doesn't mean what you are implying: once the model solved the problems and the solutions were verified, the problems weren't open any more, so a later description using the past tense is totally consistent.
Comment by traes 10 hours ago
A couple small ones that I've seen (example here [0]), but not anything of the magnitude that OpenAI and Anthropic have put out. Likely just related to token limits.
> Another question I have is whether or not OpenAI 'simply' hired capable combinatorics researchers to work on problems, and they have, and the use of the model is incidental / secondary to their work.
I think their output has reached a level that precludes this possibility, but I of course don't have any hard proof.
[0]: https://www.reddit.com/r/math/comments/1uxj3cy/after_openais...
Comment by irthomasthomas 8 hours ago
Comment by azan_ 7 hours ago
Comment by irthomasthomas 7 hours ago
Comment by kittoes 4 hours ago
I have no affiliation whatsoever with any AI company, nor any formal education outside high school, for what it's worth. Simply being curious and persistent can get you quite far in my anecdotal experience.
Comment by azan_ 7 hours ago
I don't think that comparison to p-hacking is fair. I mean not reporting price of all run is nothing like committing scientific fraud and fake results.
Comment by robinhouston 9 hours ago
Comment by antirez 8 hours ago
Comment by Chance-Device 4 hours ago
Yes. And there are many of them. I wonder what would help them come to terms with it. Seriously, people are going to be grieving over this. Loss of identity, loss of social standing, ideas of entire future lives that will now never happen. The greatest crime people may hold AI guilty of is taking away their dreams.
Comment by matteoraso 9 minutes ago
Comment by lkey 5 hours ago
Moreover, mister elite, you don't know why this press release was flagged.
I'm not sure why we should privilege your bitter speculation over more mundane possibilities.
Comment by simianwords 5 hours ago
Comment by lkey 5 hours ago
Moreover, if it wasn't flagged at all, like you suggest, then the grandparent was inventing things to be bitterly resentful about... Which is not a behavior any forum should indulge.
Comment by BigTTYGothGF 1 hour ago
Never was.
Comment by fg137 7 hours ago
Comment by antonvs 41 minutes ago
It was always mainly a website for employees of an elite.
Comment by pistoriusp 8 hours ago
Comment by defrost 7 hours ago
eg. this: [flagged] A migrant surge tests Spain's open policies (economist.com) - https://news.ycombinator.com/item?id=49131860
is clearly marked as flagged.
Unlike the current submission: Ten advances in mathematics and theoretical computer science (openai.com) which isn't [flagged].
Comment by robinhouston 7 hours ago
Comment by defrost 7 hours ago
> if you compare its rank to that of other stories with a similar age and number of points.
Ranking is complicated enough here even before weighting, speed of initial upvotes can play against ranking, number of comments and the shape of the comment tree also affect ranking. And yes, various subjects and submission sources do get weightings that impact ranking.
What's funny, to myself at least, is that any attention at all is paid to "HN front page ranking" - I've been on again off again active here since 2008 .. and can't recall ever really looking at a default HN "front page" ever.
( There's /newest /newcomments /active etc to browse and sites such as https://hckrnews.com/ )
Comment by bwfan123 1 hour ago
hah, sorry, we are plebs out here.
Comment by revetkn 7 hours ago
Comment by ltitu 1 hour ago
https://developers.openai.com/cookbook/examples/vector_datab...
How are the sales going?
Comment by halJordan 1 hour ago
Comment by 3aasgf 1 hour ago
Which is a low bar, but still.
Comment by saithound 6 hours ago
Comment by gbnwl 3 hours ago
Comment by gizmodo59 5 hours ago
Comment by curt15 7 hours ago
Comment by ianm218 3 hours ago
My guess from following this stuff quite closely is that these companies are still a couple years away from fully autonomous research staff.
[1]. https://www.anthropic.com/institute/recursive-self-improveme...
Comment by yewenjie 7 hours ago
Comment by jofzar 5 hours ago
Comment by schleck8 8 hours ago
I think we've now hit a point where 99.9% of the population gloss over these types of AI advancements because of human competence being insufficient
No human could have published this because it requires paradigm shifts (e. g. Section 5) in multiple mathematical domains. Mastering one of them to this degree is rare, mastering 3+ pretty much non existent for humans.
Comment by Chance-Device 8 hours ago
The sooner people can be broken out of their denial about all this the better, and we can start actually taking it seriously.
Comment by slashdave 2 hours ago
There is irony here
Comment by datakan 5 hours ago
Comment by Chance-Device 5 hours ago
Comment by FranzFerdiNaN 5 hours ago
I’m not a mathematician so I have zero clue what “ New upper bounds on sphere-packing density down to the Cohn–Elkies thresholds” means.
Comment by danparsonson 3 hours ago
Comment by NitpickLawyer 3 hours ago
That's not what people mean when they say "moving the goalposts". It means that people are adamant that something wasn't important/hard/impressive once the "AI" solves it. And then they come up with another thing that needs to be solved in order to prove it is important/hard/impressive. And once that happens, they do it again. And again. That's what "moving the goalposts" means.
It's also very much not a new phenomenon. It's been happening since the 1980s. As you can see from this quote from GEB by Hofstadter:
> There is a related "Theorem" about progress in AI: once some mental function is programmed, people soon cease to consider it as an essential ingredient of "real thinking". The ineluctable core of intelligence is always in that next thing which hasn't yet been programmed. This "Theorem" was first proposed to me by Larry Tesler, so I call it Tesler's Theorem: "AI is whatever hasn't been done yet."
Comment by Chance-Device 3 hours ago
Comment by ultimatefan1 7 hours ago
but it seems less likely to me than before that the types of math/science discoveries will explicitly unlock better software performance. in some sense this fits our intuitions. when top tech companies use math PhD type employees, they have them stop doing pure math research and instead focus on software engineering. these people are often very good at software engineering but not due to recent discoveries in academic mathematics, it's due to their general intelligence. to me, this is evidence that the models are getting better but does not make me think we are on the cusp of a foom style fast takeoff enabled by revolutions in frontier math (i also posted this on twitter @mlipman13)
Comment by woeirua 5 hours ago
Comment by threatofrain 5 hours ago
Comment by Ar-Curunir 5 hours ago
Like, these would be best-paper awards at many top CS conferences.
Comment by asdfologist 5 hours ago
Comment by slashdave 2 hours ago
Incredible?
> open ai announced like 15% improvement by fixing gpu kernel issue
That is... ordinary software optimization.
Comment by blovescoffee 53 minutes ago
Comment by dominotw 5 hours ago
> we are seeing frontier level math breakthroughs (ie performance that would put it in the top 100 or 1000 mathematicians in the world if it were a human, meaning top .00001% or 800/8B)
i think you have misunderstanding of what mathematicians do
Comment by DaiPlusPlus 5 hours ago
They get to make cool 3D plot visualizations of functions so obscure to me that they’re named after someone who is still alive - and/or get to work on cryptography for the NSA - I think?
Comment by kcexn 5 hours ago
It is true there hasn't been a reliable computational approach to solving these problems before. But do these proofs contribute new ideas to the mathematical corpus, or are they simply an effective method to exhaustively search the literature for the right combination of existing tools to apply to the problem?
Essentially, did these problems seem like they had an intuitive answer and were feasible to prove before, just not high enough value targets for an expert to invest time into? Or were they fundamentally difficult prior to this point and it appears that AI has done something more than just throw the problem into a big solver.
Comment by patcon 4 hours ago
Mundane incremental research is cobbled from existing citations that already appear nearby in the record.
Basically, innovative research is a measure of bridging thought and domains that were previously not bridged. It's quite concrete as a measure in the citation record.
So we can know pretty conclusively.
Puja Ohlhaver gave a talk on this[1], and ran some experiments (that I had the pleasure to support on)
Comment by Ar-Curunir 5 hours ago
Comment by QuesnayJr 2 hours ago
The sofic groups question was the outstanding question about sofic groups. Almost everyone thought that non-sofic groups existed, and there were plausible candidates, but proving a group was non-sofic was out of reach. Now that we know how to do it once, we can probably do it a lot more.
The Connes rigidity conjecture I think people thought was false, but it was a provocative claim to make. The significance of conjectures is frequently not that the answer to the question is "yes", but that we don't know how to answer the question. And now, apparently, we do.
Comment by simianwords 5 hours ago
Your worry.... is because they used the word advanced? For marketing? The word is used very appropriately here. There were PhD's who spent a big part of their career tackling these problems.
Comment by maxutility 2 hours ago
Comment by amai 1 hour ago
Comment by heaney-555 1 hour ago
Mathematicians will tear it to pieces if any of it is fake!
Comment by danielrmay 11 hours ago
> We helped prepare the manuscripts and formalize the proofs in Lean, and we take responsibility for their correctness
Offering to take responsibility for the correctness of a proof written in Lean feels like volunteering to be the fall guy in case someone finds a flaw in basic arithmetic, no?
Comment by DroneBetter 10 hours ago
Comment by traes 10 hours ago
https://leanprover.zulipchat.com/#narrow/channel/270676-lean...
Comment by jibal 9 hours ago
Comment by traes 9 hours ago
As I currently understand it, all we know is that:
- a mathematician produced a Lean-verified counterexample to the Collatz conjecture, demonstrating a bug in the kernel
- he claims that LLMs were involved somehow but pointedly refuses to specify how
- he admits that he knew about the bug before publishing the counterexample to his repository.
Perhaps not a joke (although it sure seems to me like they discovered a bug and thought falsely disproving the Collatz conjecture would be a flashy way to announce it), but at best extremely sensationalized by the above description. If you have additional context I would be happy to hear it!
Comment by danielrmay 10 hours ago
Comment by emil-lp 11 hours ago
Comment by danielrmay 10 hours ago
Comment by baq 10 hours ago
Comment by emil-lp 10 hours ago
Comment by traes 11 hours ago
Comment by DrBazza 8 hours ago
Whilst current models can't 'intuit' and come up with conjectures, they can certainly disprove some of them very quickly through the kind of grind that humans can't do. I suppose there really are some mathematicians out there today, whose last few years of study, have just been up-ended by this.
--
"Yes we are," insisted Majikthise. "We are quite definitely here as representatives of the Amalgamated Union of Philosophers, Sages, Luminaries and Other Thinking Persons, and we want this machine off, and we want it off now!"
"What's the problem?" said Lunkwill.
"I'll tell you what the problem is mate," said Majikthise, "demarcation, that's the problem!"
"We demand," yelled Vroomfondel, "that demarcation may or may not be the problem!"
"You just let the machines get on with the adding up," warned Majikthise, "and we'll take care of the eternal verities thank you very much. You want to check your legal position you do mate. Under law the Quest for Ultimate Truth is quite clearly the inalienable prerogative of your working thinkers. Any bloody machine goes and actually finds it and we're straight out of a job aren't we? I mean what's the use of our sitting up half the night arguing that there may or may not be a God if this machine only goes and gives us his bleeding phone number the next morning?"
Comment by macleginn 3 hours ago
Comment by drdrey 3 hours ago
Comment by emil-lp 11 hours ago
Comment by kingstnap 7 hours ago
Training the model is going to be amortized over other uses.
Comment by emil-lp 6 hours ago
Say that it turned out that the total cost of the proof of the Erdős unit-distance conjecture was $50 million.
Then the question really becomes: yes, these models are capable of proving important mathematical results, but at a very high cost. Is it worth it?
If a mathematician applied for a research grant of $50M USD for proving the same thing, they would have been laughed out of the bank.
What's more is that when you have a research grant, you train PhDs and postdocs, you hire new staff, and you disseminate. That is, you get much more value for the money spent.
I'm just curious what the cost is.
Comment by ianm218 3 hours ago
Comment by kingstnap 2 hours ago
Hours needed for prompt + Hours needed to check result + API costs.
You don't say "well let's add together the total yearly compensation of all the engineers and mathematicians at OpenAI that were involved" and throw that into the total cost. That's simply nonsense accounting.
The actual comparison you are making is some university researcher weighing between getting a grad student (several tens of thousands of dollars) vs typing up a prompt and sending a request to OpenAI for inference (as mentioned in the article, around $2000 in API and maybe a few hours for the prompt and harness).
Comment by z7 10 hours ago
Comment by traes 10 hours ago
Comment by traes 10 hours ago
Comment by avaer 10 hours ago
Comment by traes 10 hours ago
Edit: Oh, are you suggesting they just use it to privately improve their models? I imagine a few more correct proofs would have a very marginal benefit, if any. Also, they'll probably just get extracted, meaning it still gets out but OpenAI doesn't get to fancily announce it themselves.
Comment by asdewqqwer 9 hours ago
Comment by simianwords 10 hours ago
Comment by lifeisstillgood 10 hours ago
Which is very suggestive - if after everything they are not fully loaded then the next gazillion data centres being built look unlikely to be needed.
Comment by paxys 1 hour ago
Comment by lwansbrough 9 hours ago
Comment by traes 10 hours ago
Comment by Davidzheng 7 hours ago
Comment by simianwords 5 hours ago
If anything it proves more data centres are needed. That's literally the only reasonable conclusion from this news.
Comment by lifeisstillgood 1 hour ago
Are you arguing there is not an AI bubble, and that all the DC buildout is fine, going to be profitable etc?
I am not looking for a online slanging match - just looking for a different point of view
Comment by simianwords 1 hour ago
I’m not participating in the slinging match but it’s very very weird that you think it’s some established thing that these companies won’t make profit. A lot of hubris must go in this kind of thought. Like.. do you all think everyone’s playing musical chairs?
Comment by svieira 1 hour ago
Comment by simianwords 1 hour ago
Comment by 0x5FC3 11 hours ago
Comment by traes 10 hours ago
Comment by 0x5FC3 10 hours ago
Comment by jryle70 3 hours ago
Comment by simianwords 10 hours ago
Comment by 0x5FC3 10 hours ago
Comment by frozenseven 10 hours ago
Comment by shimman 17 minutes ago
These companies desperately want a return to serfdom. If they didn't come off as so anti-human the public backlash wouldn't be so great.
Comment by jacki 10 hours ago
Comment by jgeralnik 9 hours ago
This was not a problem that was for sale
Comment by heaney-555 1 hour ago
Comment by piker 11 hours ago
Now I understand that it’s mostly the super stars benefitting from the increased attention. Folks who are less established don’t share in that glory. But on the other hand it seems like an exciting time to go even deeper for in various specialties of math by deciding where to focus these powerful tools. For every conjecture defeated some seven or eight new ideas open up. Our path through that combination will be set by creative and curious human mathematicians.
[edit: deleted a distracting comparison to Chess]
Comment by energy123 10 hours ago
Comment by traes 10 hours ago
Comment by FranzFerdiNaN 5 hours ago
Which is less interesting work. And you probably need to do the hard grunt work by hand first to develop the skills and intuition to be able to verify an AI-generated result. So you can’t outsource everything to AI without loss of skill.
Comment by dash2 2 hours ago
Comment by aabhay 11 hours ago
Comment by kzrdude 6 hours ago
Thinking of this a little bit with the perspective of every new proof as a burden, dumped for review by actual mathematicians.
Comment by traes 11 hours ago
Comment by artninja1988 5 hours ago
Was this different before chess computers were invented?
Comment by anematode 10 hours ago
Comment by energy123 9 hours ago
Comment by sashank_1509 5 hours ago
Comment by ratmice 10 hours ago
Comment by traes 10 hours ago
Comment by ratmice 9 hours ago
Comment by traes 8 hours ago
Comment by jibal 10 hours ago
If you translate that to math, then all you get is math competitions, not math as a career. Of course the translation isn't nearly exact ... there's a lot more room for professional mathematicians because the math space is far more vast than the chess space and can't generally be cranked out mechanically (we have proof).
P.S. The response is nonsense ... I explained exactly why it's awful (others have too) and the response doesn't in any way refute the explanation ... rather it offers up a ridiculous strawman.
Comment by piker 10 hours ago
Comment by baq 10 hours ago
The most disappointed folks will simply drop out, but the enthusiastic ones will keep going and with luck make up for the ones who decided to quit. Chess and go certainly went this way.
Comment by traes 10 hours ago
Comment by artninja1988 7 hours ago
Comment by laichzeit0 6 hours ago
Comment by slashdave 2 hours ago
Comment by slashdave 2 hours ago
Comment by Davidzheng 7 hours ago
Comment by artninja1988 7 hours ago
Comment by christofosho 7 hours ago
I'm sure they must do some of this type of work, right?
Comment by braneloop 6 hours ago
Comment by amazingamazing 6 hours ago
Comment by throwaway198846 3 hours ago
Comment by adroitboss 5 hours ago
Comment by adroitboss 5 hours ago
Comment by beering 3 hours ago
Comment by slashdave 2 hours ago
We already know how to solve all of these issues. What we lack is collective political will.
Comment by solenoid0937 5 hours ago
Comment by miltonlost 1 hour ago
Comment by ltitu 1 hour ago
Comment by frenzyguy 6 hours ago
However, I was looking at the proofs and reason explanation and openAI should be more explicit in how the work has flown. I find the models have jumped hoops in some places of the proofs, that can be hard to track. In fact, when a paper is published you usually get a review and if no reviewer understands they ask you to further explain the thought process. It will be fun to see if this happens here.
Comment by scuppernong 2 hours ago
Comment by bifftastic 7 hours ago
Comment by QuesnayJr 6 hours ago
Comment by kingstnap 8 hours ago
And yet this is the exact same company that has screwed up their android app so bad that the latex N^3 rendering problem makes it so having it explain it to me crashes the app.
Truly jagged beyond belief.
Comment by melagonster 8 hours ago
Comment by xyzsparetimexyz 8 hours ago
Comment by woeirua 6 hours ago
Comment by s_Hogg 10 hours ago
Comment by defrost 10 hours ago
Comment by amazingamazing 6 hours ago
Comment by readthenotes1 9 hours ago
Comment by sashank_1509 4 hours ago
They can’t even drive me anywhere I want, though I suppose they’re getting there. I don’t think LLM companies should be surprised when rest of society hates them. They’re literally bringing in a dystopian WallE like society where most demand for human work is destroyed.
Comment by matteoraso 1 minute ago
Comment by unknownian 3 hours ago
Comment by eadwu 2 hours ago
Stop coping and deluding yourself mate.
To begin with, whether AI is the one doing the discovering or not makes no difference. Any "pure math" person would aim to understand regardless - and would be quite glad that they have a longer paved path.
Any mathematician in academic or industry is more than likely not a "pure math" person (tainted by capitalism).
Comment by unknownian 21 minutes ago
>Any mathematician in academic or industry is more than likely not a "pure math" person (tainted by capitalism)
Ignoring that I meant pure as in non applied math, let's just make it clear: you agree that mathematicians who are against capitalism encroaching on this process should be allowed to dislike it without criticism of being pretentious?
Comment by deyiao 10 hours ago
Comment by k2xl 10 hours ago
Comment by utopiah 10 hours ago
Comment by utopiah 7 hours ago
Comment by luciana1u 10 hours ago
Comment by baq 10 hours ago
Comment by xyzsparetimexyz 8 hours ago
Comment by utopiah 7 hours ago
Comment by foobar10000 6 hours ago
The non-sofic group one is definitely a big deal - would have been a Fields medal if discovered by a human.
Comment by zkmon 10 hours ago
AI has no self-awareness. It's a tool. When you assemble a furniture using a screw driver, the torque force interacts with the molecular forces inside the metal and miraculously it transfers the force to the screw though a clever geometry design, communicating the force to the screw to turn it in a certain way.
Do you attribute the build to the tool? The "system's contribution" is helped by many other things all the way down to chips, datacenters and power generation. If the authorship requires attributing to a tool, then it should happen all the way down.
Comment by raincole 10 hours ago
A slightly smarter highschooler could write these. I could write these. It's clear as day that the LLM, not the human, did the heavy lift. It'd be ridiculous to give full credit to whoever wrote the prompt.
[0]: Not one of the proofs in the linked article, but from OpenAI too.
Comment by ben_w 10 hours ago
I think you're over-estimating what a smarter highschooler could write.
A "finite loopless undirected multigraph" could have been explained to me at that age if we'd taken Discrete rather than Mechanics and Pure (and one module of Stats) in my two A-levels* in maths and further maths; but from what I saw of the Discrete module, neither:
Every finite loopless multigraph with no bridge possesses a cycle double cover, without additional assumptions such as cubicity, planarity, connectivity, or higher edge-connectivity.
nor: repeated-edge closed trails masquerading as cycles
would have been something we'd have learned. But more importantly, we absolutely didn't have a feel for how much effort one needs to put into making sure the proof is right, so if one of us had been hypothetically asked to write a prompt it would've been no more than half that length, and missed most of the bullet points.* For those not from the UK: A-levels are between secondary school and university, when aged 16-18. Functionally they are university entrance qualifications: https://en.wikipedia.org/wiki/A-level_(United_Kingdom)
Comment by yaqubroli 10 hours ago
Provability is just going the way of computation. John Napier had to manually compute logarithm tables over decades and was recognised for his work; now that same work could be performed by a 10 year old with a calculator in an evening.
Comment by esikich 10 hours ago
Comment by ipnon 9 hours ago
Comment by ascots 48 minutes ago
Comment by raincole 9 hours ago
[0]: e.g. "go through wikipedia's unsolved math problem list and solve them".
Comment by zkmon 10 hours ago
Comment by raincole 10 hours ago
Comment by zkmon 9 hours ago
Comment by Anon1096 8 hours ago
Comment by mathisfun123 10 hours ago
> In particular, proofs for special graph classes, constructions of cycle covers with some edges covered other than twice, bounded-length or prescribed-cycle variants, reductions to another unproved conjecture, computational verification through any fixed graph size, and candidate counterexamples without a complete nonexistence certificate are insufficient.
which is infact a very important part of the prompt.
Comment by don_esteban 8 hours ago
Comment by skinner_ 3 hours ago
Take a grad student with a perfectly good understanding of what a proof is. Their supervisor gives them a major problem to work on. Almost always, the problem is too hard, the student comes back with partial results, and student and the supervisor iterate from there. Now imagine that they have an unusually cruel and unreasonable advisor who tells them, do not dare to talk to me until you've fully solved the problem. This paragraph is exactly that. It's there exactly because the underlying system is smart enough to know that real mathematicians do not work like that.
Comment by famouswaffles 4 hours ago
To name a few:
- https://xcancel.com/DmitryRybin1/status/2079904005652893709
- https://archive.ph/2w4fi (https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...)
Comment by esikich 10 hours ago
Comment by NitpickLawyer 10 hours ago
Comment by cure_42 10 hours ago
Comment by dgellow 10 hours ago
Comment by NitpickLawyer 10 hours ago
Comment by ben_w 10 hours ago
When the tool is a 3D printer, or any CNC system really, you bet I attribute a build to it.
I could also attribute the operator; there is no contradiction, it's a free choice, just like saying "I am in Berlin" does not contradict "I am in Germany".
Comment by naasking 10 hours ago
What is your mechanistic model of self awareness that yields this conclusion?
> It's a tool
Does your model suggest that tools can't have self awareness?
Comment by Delk 9 hours ago
A language model (or an image model or whatever) cannot even be sentient, and I think sentience is a prerequisite for awareness.
Even if we express a lot of our subjective experience with words, the language is just a symbolic representation of those experiences. The qualia themselves, even those that are quite abstract, are rooted in our physical presence and evolution.
You can't have an understanding of what hunger or physical pain feel like if you have no need for food or a sensory capacity for feeling pain. You can't understand what loneliness or pride at an achievement mean if you don't have a neural network wired to value social connection or status. We value connection because we're social animals that have needed each other for survival.
Even the more abstract of our subjective experiences are in some way rooted in our physical evolution.
I see no reason to believe that a neural network built entirely based on the symbolic level of language could have the features needed for the subjective experience itself.
AI awareness might actually be more believable if that awareness manifested itself in an entirely different way than in humans. But if we assume awareness because outputs resemble what we consider meaningful as humans, yet the neural network has had no inputs or evolution that could form the actual basis of human-like experience, I think we're seeing something that isn't actually there.
Comment by naasking 2 hours ago
There is no objective evidence of qualia. All evidence of qualia are vocal or other expressions of belief in qualia. Perceptions clearly exist and are observable, subjective experience and qualia, not so much.
> I see no reason to believe that a neural network built entirely based on the symbolic level of language could have the features needed for the subjective experience itself.
If your objection is to models based on "symbolic level of language" which you think lack semantic understanding of, say, trees, you should ask yourself how our brain, based on physics which also lacks any semantic category for trees, can somehow develop a semantic understanding of trees. All of these appeals to differences with the brain never seem to acknowledge that fundamentally, the brain has the same explanatory gap with physics.
> But if we assume awareness because outputs resemble what we consider meaningful as humans, yet the neural network has had no inputs or evolution that could form the actual basis of human-like experience
This assumes a lot. It seems very possible to me that intelligence inherently develops a map of natural categories (natural kinds), and language naturally develops around such categorical understanding. Semantics are then fundamentally the network of associations between categories, eg. there is no fundamental difference between symbols and semantics, and the latter cam be inferred from the former, and that's exactly what LLMs do, and why the semantic maps between different languages are so similar and how they can translate between languages.
Comment by woeirua 6 hours ago
Comment by Delk 3 hours ago
You didn't address any of what I wrote, let alone provide any counterarguments. Which part of what I wrote do you think was wrong?
Comment by naasking 2 hours ago
For example, the Turing machines and the lambda calculus don't look anything alike, but they are fundamentally interconvertible, and so in a real sense they are fundamentally equivalent. Without a model, all of your arguments are completely unconvincing for exactly the same reasons, eg. that there may exist many paths to fundamentally equivalent ends.
Comment by Delk 18 minutes ago
Hunger as a concept doesn't mean anything without the physical need. Politeness or bluntness, even in writing, don't mean anything without social dynamics. And we have social dynamics (and neural structures that directly process social cues and associated feelings) because we've evolved into social animals for whose survival that was important.
I see no reason to believe that a model trained only with symbolic representations, with no connection to the physical world phenomena that those symbols represent, could contain the subjective experience itself.
Neural network models may be able to derive novel (or at least novel-looking) output rather than just an obvious rehash of their input, but I don't think any set of bytes can fundamentally contain information that was never entered into it. (Even if e.g. a model produces previously unknown mathematical results, those results can in principle be derived from the information that they were trained with.)
I'm not saying that artificial neural networks couldn't, in principle, be aware. ANNs and biological neural nets may be equivalent in the sense that any information and processing structures represented by a biological one could in principle be represented by an artificial one. If that's the case, and awareness is purely a product of our neural systems as materialism would imply, it should be possible for an ANN to be aware, too.
But when the model has been trained with only language, and IMO the subjective experience can't be derived from the symbolic representation alone, I can't see how the model could include the actual subjective human experience.
An AI model could of course have an awareness and subjective experiences that are totally different than our human experience. But then the fact that it happens to produce output resembling what humans find meaningful shouldn't be considered indicative of such awareness.
This is of course more of a philosophical argument than a technical one, and I'm happy to hear counterarguments, but not on the level of off-hand dismissal.
Comment by perching_aix 10 hours ago
(*) Even if we hack around this and just do the usual trick of simply laundering statefulness to a higher level, in this case the context window being fed in, I fail to identify (**) a representation of its own state in these bodies of text that it'd be meticulously maintaining. I further fail to identify how it could be hidden or maintained, considering I control like half of it. The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either. Difference being, I actually have one (I'm "on-line").
You'll sometimes catch models mixing up who's who and how many who-s there even are for example.
(**) I did wish for something hidden though, so maybe it's just concealed? The same way people can encode a lot more of their emotional and mental state than normal into text if they read and write a lot of it, I'm aware of research that suggested the same for LLMs, albeit I cannot cite it. Maybe those phrasing signatures are just alien to me and will never pop out. Either way, I'd expect researchers to stumble upon this during interpretability studies, and either they haven't, they have but it wasn't popsci adopted, or they're keeping awfully tight lipped about it. If you know of anything like this, your turn now, would be happy to learn.
I do wonder how reasonable it is to expect e.g. a single maintained identity though. Maybe it isn't?
(*) Another way to hack around this of course is to just precompute some internal "self-awareness states" and hop around between them. Probably the closest to what the models are actually "doing".
Comment by ben_w 9 hours ago
> a hidden representation of self that is continually tended to
This sounds like a personality? They act like they have one of those. It may be an illusion, and even if it isn't an illusion it is unlikely to be anything like the source (us), but they act like it.
> I further fail to identify how it could be hidden or maintained, considering I control like half of it.
Indeed you control everything about a local model, and much of the context of even a remote model. But the state of activations and circuits in SotA AI is hidden in similar ways to those of synapses in your head: difficult to decipher even with probes monitoring the signals directly, and often not emitted at the normal output.
> The best you could ascribe it is a meticulous maintenance of a persona the user is talking to, but then that doesn't necessarily represent the model's internal state, the same way my own words here aren't doing so either. Difference being, I actually have one (I'm "on-line").
While we can be confident that LLMs make up personas etc., it is insufficient to go from "that doesn't necessarily represent the model's internal state" to "therefore it doesn't have one".
> You'll sometimes catch models mixing up who's who and how many who-s there even are for example.
I've, unfortunately, also experienced this with humans. Perhaps they were losing their self-awareness at the time? I do wonder if old-age dementia does that by the end, though the person in question didn't ever get diagnosed with that.
> If you know of anything like this, your turn now, would be happy to learn.
Do you mean like these, or something else?
• https://researchportal.hkust.edu.hk/en/publications/decoding...
Comment by perching_aix 7 hours ago
Not quite what I meant, but it's also not entirely unrelated I guess? Personality to me is like a natural bias. It does also shift over time, and is also an internal bit of state. I guess in some respects it can also be self-referential, like personal convictions.
> Perhaps they were losing their self-awareness at the time?
I do think it is entirely possible for people's self-awareness to shift, yes. Or more precisely, I do model things that way.
> Do you mean like these, or something else?
They're adjacent, but I more meant something like these:
https://arxiv.org/abs/2410.03768
https://arxiv.org/abs/2310.18512
https://arxiv.org/abs/2605.26537
So basically, steganography. The difference is that these papers investigate from the perspective of separate LLM instances covertly exchanging information between each other. This is in contrast with the scenario I'm laying out, where an LLM's past state is exchanging information with its future state, continuously representing and modulating a concealed internal state of some sort. And then that state just so happening to be some sort of self-referential meta state.
And the best inkling I have towards this is basically: https://www.youtube.com/shorts/WP5_XJY_P0Q
But then I don't think there's enough covert channel bandwidth in the agent replies for anything interesting like this.
Comment by naasking 2 hours ago
I don't see why an LLM could not have a sense of identity or personality while it's evaluating a specific prompt, or even change self awareness while evaluating a prompt since many outputs model a back and forth conversation. My point is that without a mechanistic model of what "self awareness" means, we have no way of truly evaluating such questions, we're just hand waving vague intuitions about what it could mean.
Comment by petilon 1 hour ago
Sam Altman has said "If superintelligence can't discover novel physics, I don't think it's a superintelligence." Is that the test? How far away are we from AI discovering novel physics? It seems within reach.
Comment by antonvs 36 minutes ago
Comment by petilon 30 minutes ago