OpenAI’s amazing — but vastly oversold — new model Astra

Posted by champagnepapi 20 hours ago

Counter25Comment9OpenOriginal

Comments

Comment by keeganpoppen 17 hours ago

people who throw out “x is a fallacy” very seldom have anything interesting to say, and this point is no different. the thing is not released yet. there does not exist any easy way to communicate how strong any particular llm model is en toto, so people inevitably look for stories, metaphors, and the like to help explain the story. saying that it is “vastly oversold” is comically absurd for a model that has not been released yet, especially through the lens of a completely unfalsifiable framework for contextualizing it. here, i’ve got a “fallacy” for you: this whole post reeks of “no true scotsman”: it is bold to claim that the model is oversold in its abilities when it has racked up this many novel proofs before even being released widely, but hiding behind “that doesn’t mean it is good at ‘math, generally’” is absolute weasel language— it invites proof-by-example in a way more flagrant and devastating than anything the author points out about the discourse around the model itself.

Comment by malshe 15 hours ago

I wonder if it can write like a normal human being. In my experience the writing is becoming worse the more advanced a model is. Anything written by Fable is practically unreadable.

Comment by axus 15 hours ago

I like how they use SAT Math and Verbal sections for an analogy, and talk about how great OpenAI is at some kinds of math, but leave out how it's even better at written/verbal language.

It wouldn't hurt their argument to observe how well LLMs do at reading and writing, but it's probably painful for them to admit.

Comment by kelseyfrog 11 hours ago

Anyone who thinks that making advances in mathematics implies competency in other aspects of life has never been married to a mathematician.

Comment by igor47 19 hours ago

Can someone give me the current consensus on Howard Gartner/multiple intelligence vs. g, or general intelligence? Is this even something that people still discuss and research in academia, or did the whole field get tainted by accusations of racism and counter accusations of censorship? Gary's claim here rests on an implicit disregard of general intelligence, which seems counter intuitive to me, but it's been a long time since I've looked into it

Comment by SpicyLemonZest 18 hours ago

I don't think his claim is really related to general intelligence as such. What he claims is that the intelligence of the underlying models, general or otherwise, has not advanced as far as is commonly believed. He thinks that the observed improvement is actually attributable to verification tools, so it won't generalize to problem spaces which aren't verifiable enough to build such tools.

Comment by perching_aix 19 hours ago

I envy the folks who have the energy to speculate this much about an unreleased product/service, and spend this kind of - human - reasoning effort reflecting on blatantly worthless internet posts.

You really don't need to break open the fallacy dictionary to see why those tweets are phony, or to telegraph Astra as just an incremental [0] improvement that's even better tuned for math than what came before it. It's the obvious direction of development.

[0] There's a mathematician guy I follow on YouTube who keeps taking LLMs for a spin, and the primary failure mode seems to be persisting. It's not dissimilar to any other field; the models are chatterboxes, and keep going off about stuff that doesn't matter, while quickly jumping over things that do. They're also comparatively slow. According to another mathematician's review of the 250 page paper OAI put out of those 10 breakthroughs, the former persists with Astra.

I wonder if Astra can run on those Cerebras wafers. An order of magnitude faster inference would at least make the iteration process quicker. But then they were announced for Sol too, and they're nowhere to be found. The 2.5x fast mode is nice, but it's a far cry from the 750 tok/sec suggested with Cerebras.

Comment by dude250711 19 hours ago

Whatever, as long as it does not burn Codex tokens too fast.

Comment by eec33 17 hours ago

[dead]

Comment by semiquaver 18 hours ago

The cope is tangible. Article is drenched in flop sweat to a remarkable extent, even for Mr. Marcus. I cannot see the goal posts, they have been moved so much.