Grok 4.7
Posted by meetpateltech 1 day ago
Comments
Comment by moojacob 1 day ago
Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.
However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.
My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
Comment by imron 1 day ago
Grok has its own feel too. It's not as bad as Claude, but one of the things that bugs me is that it is far too terse.
It regularly seems to come up with terms and descriptions for things in its chain of reasoning and then uses these terms in its output assuming you understand what it's talking about.
I find I often have to ask it to re-explain what it means.
Comment by runeks 14 hours ago
GPT does this all the time, too (both Sol and Astra). I constantly have to tell it to not use terms that were not part of the initial prompt.
Comment by imron 14 minutes ago
Hah, yeah even when you put it in AGENTS.md or a skill.. constantly having to remind it.. "what does AGENTS.md" say about doing that?".. Thinking.. Thinking.. "Oh, it says I should never do that, I'll remember that next time.."
Next session - same thing.
Comment by glub 13 hours ago
Comment by ls612 8 hours ago
Comment by taspeotis 23 hours ago
Separately have been using Grok 4.6 for a bit and it's also pretty concise.
Comment by iamflimflam1 1 day ago
Comment by bilbo-b-baggins 20 hours ago
I’m pretty sure the big bois don’t do it because it would undermine “confidence”.
Seeing a model output “Oh I should just delete blah. Wait blah is a production service, I shouldn’t touch that. Maybe I can gain access to blah? Oh the aws cli isn’t signed in to blah. I see kubectl has access to blah though! Wait, I should ask user permission first.”
Yeaaaaah. Thinking tokens are fuckin’ wild.
Comment by demibabs 17 hours ago
Comment by geetee 17 hours ago
Comment by tyre 18 hours ago
Comment by glub 13 hours ago
Could be something very stupid like - "I don't have ffmpeg available here. Should I install it? No, I can't. I'll proceed doing something that will take me 100x more tokens and wall clock just to avoid adding a dependency." I can then just stop and say - you've got nix flake there, just add it.
That's impossible with western models. The only way is to ask why it did something stupid when it already spent 50% of your weekly quota and produced millions lines of slop.
Comment by imron 18 hours ago
Comment by tempay 18 hours ago
Comment by greenavocado 19 hours ago
Comment by Zambyte 1 day ago
Comment by imron 23 hours ago
Comment by octoberfranklin 19 hours ago
I suspect that they specifically train Grok to be able to work well with military personnel -- speaking the way they speak: brief, to the point, efficient communication. Personally I really like this. Claude sounds like some demented clown from the marketing department.
Comment by imron 18 hours ago
Comment by shuwix 17 hours ago
Comment by JimDabell 20 hours ago
I’ve noticed Astra doing this a lot as well.
Comment by caminante 1 day ago
Anecdotally, I have noticed the same in the past week. It might just be anecdotal or driven by a long context window.
Comment by suburban_strike 5 hours ago
Comment by smashers1114 1 day ago
Comment by a2dam 1 day ago
First, after a while it's just as grating as Claudeish. Second, my hunch is that it constricts the actual thinking of the LLM, like the same way that Newspeak does in 1984. It shrinks the range of thought that can be expressed if used as an input.
I think the real way to do it is to have another Claude entirely deal with the user as a liaison, but to keep the thinking in whatever format it came in.
Latent space reasoning, if you think about it, is exactly this to a crazy degree: why even formulate a thought as words if you can just keep it as matmuls until the user needs it? And then, if the user needs it, have it always specifically formulated for the user by another LLM rather than constrict its range of thought? Anyway, that's my take.
Comment by smashers1114 1 day ago
I do think an infrastructure where another Claude retranslates the output would be better. Oftentimes I forget to put it in the actual prompt and when I receive back 8 paragraphs of Claudeish I ask for it then.
I would have to disagree that it gets as grating as Claudeish though. Its just direct and professional instead of ring-around-the-rosy clickbait.
Comment by powvans 22 hours ago
“I would have to disagree that it gets as grating as Claudeish though.”
It’s hard to imagine anything more grating than Claudeish. To quote Rainer Wolfcastle, "My eyes! The goggles do nothing!"
Comment by mrandish 19 hours ago
I've found that prompting any constraint on output (length, style, vocab, even simple formatting) not only places additional cognitive load on the model, which burns some of whatever cognitive budget is available, it will also often skew the output in other subtle and completely unrelated ways.
Since I found this artifact interesting, I did some pretty extensive experiments a couple months ago. The increased load is real, although it may not be apparent if you're not near any cognitive boundaries. The subtle skew, however, seems nearly ever-present regardless of load.
Comment by evulhotdog 17 hours ago
Comment by mrandish 5 hours ago
While extensive, my tests were just following my curiousity, not controlled, exhaustive or well-documented. I identified about a dozen prior sessions of varying length and complexity to test and downloaded them with a browser add-on. I then removed all other user prompt instructions except for the formatting instruction. A test would typically involve changing the wording of the formatting instruction ranging from brutally simple to detailed and complete, then starting a new session, seeding one of the test sessions and continuing it. To get a feel for baseline inter-session variation, I also tried running the exact same prompt/session multiple times back-to-back, at different times and on different days of the week.
Once I identified a promising prompt candidate, I'd make it the formatting instruction in my regular, daily-use prompt for a few days. I quickly got a feel for how seemingly minor user prompt variations impact response quality, compliance and tone across fresh sessions as well as those in various states of context rot, drift, decay and cliff (<--my nicknames for the distinct flavors of session degradation, not technical terms).
My overall conclusion was that every instruction, no matter how minor or unrelated it seems, has some, real impact on the model's cog load, attentional focus and/or attentional weight budget. Both how these impacts manifest and what causes more or less impact is often extremely counteriintuitive. To more fully understand this, I eventually, got to the point of testing null case variants, such as the entire user prompt being one sentence completely unrelated to text formatting or the session topic, like: "Don't reference the cartoon character SnagglePuss" (in a deep dive on ancient Sumerian clay tokens). Similarly, a simple one sentence prompt requesting something the model already always does naturally also has a cost (eg "Capitalize proper nouns"). As others have observed, heavy emphasis, absolute prohibitions or emotional weight in prompts also tend to have outsized impact in both skew (impacting unrelated output tone/style) and in accelerating session degradation. "Avoid referencing SnagglePuss when you can" would have equal compliance but fewer downside impacts than "NEVER reference the cartoon character SnagglePuss" in sessions starting to degrade.
There were also surprises, such as when I was scanning transcripts of an older, longer session and noticed the LLM was doing number formatting almost perfectly. On looking at the active user prompt at the time (I keep a log of every user prompt change I make for every model), it didn't even reference formatting at all. More experimentation showed it a result of the LLM gradually mirroring my consistent use of formatting structure in my prompts over a long session (in which I never mentioned anything about formatting). Unfortunately, that mirrored trait doesn't persist to new sessions and reaching that point requires a substantial number of rounds burning quite a bit of context window.
After spending time surfacing the impacts of just changing lightweight user prompts so they could be observed (which are the lowest priority prompts a model gets), I now wonder just how much more 'brilliant' the models we use daily would be if they didn't have dozens of pages high-priority manufacturer prohibition prompts we never even see weighing them down. We've only ever seen these frontier 'racehorses' when they're already pulling a heavy invisible wagon.
Comment by TuxMark5 1 day ago
Comment by nomel 1 day ago
I don't think this is true.
They have to express themselves as tokens. The meaning of those tokens doesn't have to be text. See any model that can handle images/video. Also, I don't think math, svg, etc, are "natural" language.
And, only the final expression is tokens. The intermediate layers, with the encoded concepts, aren't "natural language".
But, to address your concern (which nobody can disagree with, since even humans can't fully express through text/pictures), potentially: https://news.ycombinator.com/item?id=49758615
Comment by _puk 1 day ago
The model isn't limited to concepts that can be expressed in natural language.
It's only once the AI gets to the output layers that natural language comes back into play.
After all, they're all made out of weights[0].
Comment by lelanthran 16 hours ago
How do we know for sure? We don't even know how the emergent properties we see actually emerged?
For humans we know for sure that people sometimes have concepts that they have no word for (the reason the phrase "It's on the tip of my tongue" is a phrase, after all).
We don't know this for LLMs. When it makes new phrases, it's always a mixup of two existing words hyphenated (aside, that also seems to be the limits of SOTA models creativity - join two unrelated words together with a hyphen).
LLMs never respond with "It's on the tip of my tongue" type responses, indicating it has a concept but cannot remember (or does not have) a word for that concept. Every human, pre-speech-age, has managed to express or convey concepts that they had no word for.
So, no. I'd need a citation, preferably multiple, that did the trials and found that a model can generate concepts for which it does not have any words for.
Comment by helloplanets 17 hours ago
Even if the input is in plain English, the model never sees any words, tokens or glyphs to begin with. It's vectors all the way down.
Comment by timacles 1 day ago
Comment by nomel 1 day ago
Comment by Dylan16807 1 day ago
Comment by booty 1 day ago
1. Is natural language holding LLMs back by some %? 2. Is natural language serving as a hard gate that will prevent LLM intelligent progressing past some specific point?
The answer to 1 seems like an obvious yes to me.
Your thesis says the answer to 2 is "yes." That doesn't feel right to me. Think about all of the humans who have pushed various fields forward: Einstein, Newtown, Bach, whoever. If natural language doesn't prevent an entity from surpassing humans in one intellectual field, why would it prevent an entity from surpassing humans in all intellectual fields?
(To be clear, I'm not claiming superintelligence will or won't be achieved; I'm considering your specific thesis about whether or not natural language will be a hard gate)
Comment by indigo945 1 day ago
By the way, how good is Claude's Hopi?
Comment by dist-epoch 1 day ago
Comment by faeyanpiraat 1 day ago
Comment by StilesCrisis 1 day ago
Comment by cjonas 20 hours ago
Comment by oxidant 1 day ago
Comment by _boffin_ 1 day ago
Comment by bel8 1 day ago
Comment by tempest_ 1 day ago
It burns more tokens but is the only way to get tolerable text.
Comment by hungryhobbit 1 day ago
Comment by tempest_ 1 day ago
https://code.claude.com/docs/en/hooks-guide#agent-based-hook...
Comment by hungryhobbit 1 day ago
Comment by LPisGood 1 day ago
Comment by kekebo 1 day ago
Comment by MisterMunchkin 1 day ago
Comment by ffsm8 1 day ago
Literally every one, even 1-2 prompts later it starts to go back
Comment by neomantra 19 hours ago
It’s been really productive and I’ve been asking my agents to communicate using it more and more. I believe it’s relieved my cognitive load a bit while working with them.
Comment by SoMomentary 1 day ago
Comment by el_benhameen 1 day ago
Comment by junon 1 day ago
Comment by snapplebobapple 1 day ago
Comment by faangguyindia 12 hours ago
Comment by BatteryMountain 1 day ago
Comment by BatteryMountain 15 hours ago
Comment by fr2029 14 hours ago
Comment by guluarte 1 day ago
Comment by jasonjmcghee 1 day ago
That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
Comment by vessenes 1 day ago
Comment by boc 1 day ago
Comment by jitl 1 day ago
Comment by boc 3 hours ago
Comment by vintermann 1 day ago
Comment by svachalek 1 day ago
Comment by vintermann 1 day ago
Comment by Lucasoato 1 day ago
I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?
Comment by TomGarden 1 day ago
The weird thing is, that's not what AI models seem to be doing. The prose is just weird.
Comment by unshavedyak 1 day ago
This happens most though when the speaker doesn't (or care to) understand their audience.
Eg i find effective communication requires expertise in both the subject matter domain but also the reference of the listener. Eg in ELI5 framing, if you don't know what information 5yr olds are expected to know you'll do a poor job at an ELI5.
It often feels like Claude does poorly at both framing the response relative to what it "thinks" the listener knows, but also the prose is... sideways, just weird as you said.
Comment by TomGarden 1 day ago
If I don't grok an elaborate explanation, I can ask for clarification. If it's explained to me in an overly simplistic or unnuanced way, I'll walk away with a false sense of understanding.
That said, I'm sure we all have very different concentrations of these types of people and problems around us. I've definitely met some engineers who seem to actively try to make their language incomprehensible
Comment by pixl97 1 day ago
Comment by cyanydeez 1 day ago
It is unsurprising that a LLM fails, without coaching, to effectively communicate.
Comment by fearmerchant 1 day ago
Agreed. Do you think it's due to that EU issue of making AI text be identifiable?
Comment by flipthefrog 1 day ago
Comment by MisterMunchkin 1 day ago
Comment by superjan 1 day ago
I should try adding these tips to my system prompt. Is there a shorthand to describe such language use? I am not a native English speaker.
Comment by svachalek 1 day ago
As for the wording of the prompt, you're pretty on point, I created a custom output style targeting mostly the first two you have there. Some people have wording that demands a certain technical standard or uses fancy words to describe what to avoid, but I haven't seen evidence those work better than asking plainly and I suspect the opposite: LLMs mimic the user to a degree so talking to it in terms of technical specifications and fancy words is an invitation to get them back.
Comment by samuelknight 1 day ago
Comment by Aperocky 1 day ago
When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.
Comment by fragmede 1 day ago
Just because something is difficult to understand doesn't mean it's fraud, although if someone is trying to dazzle you with clever words and names of institutions you recognize because they are selling you something, there's a good chance they're lying to you in order to get some money from you.
Comment by Pannoniae 1 day ago
Sure the explanation will oversimplify a lot but then you can expand it recursively if needed, you gotta start somewhere.
Comment by Aperocky 17 hours ago
If the presenter can't divide a problem until it reaches a series of independently simple concepts, then there's usually something fishy going on.
Comment by menaerus 14 hours ago
Comment by Aperocky 9 hours ago
Comment by includenotfound 1 day ago
You just simplified most of the problems people work on down to cancer complexity. Ironic, isn't it?
That's also simply not the case, most people are building CRUD apps with some frontend code and some accessory stuff like build systems etc., which while complex, can still be expressed in very plain, easy to understand language for anyone who's a bit technical.
Does not excuse the Claude slop.
Comment by fragmede 1 day ago
Solving the problem right in front of you is easy. Stepping back and asking: is that a problem to be solved, is infinitely harder.
I did not use Claude to write my comment, so I don't know where that is coming from.
Comment by includenotfound 1 day ago
Comment by r_lee 13 hours ago
this is specifically an Anthropic problem, maybe due to their heavy use of Claude to train Claude itself?
Comment by thesmtsolver2 1 day ago
Part of intelligence is knowing your audience and communicating efficiently.
Comment by cruffle_duffle 1 day ago
Bingo! And on this axis many SOTA models fail miserably. These things are acting on my behalf under my direction. All the supposed intelligence in the world means fuck-all if nobody can understand it.
And like somebody else said… when meat-based humans talk like Claude does, it almost always means they either don’t understand what they are talking about, or are actively trying to conceal something and are a fraud. Not always, but almost always.
Comment by yread 1 day ago
Comment by grababner 1 day ago
Comment by michaelmrose 1 day ago
Comment by WarmWash 1 day ago
Comment by moojacob 1 day ago
I am a huge fan of Gemini Pro for chat... gemini somehow just knows the most obscure stuff. I'll double check something Gemini said and find the source is deep inside a hard to access scientific paper. Google just has the best index of the internet.
Comment by haellsigh 1 day ago
Comment by StilesCrisis 1 day ago
Comment by r_lee 13 hours ago
Comment by WarmWash 1 day ago
It's best for brain storming, rabbit holes, and image recognition.
Let the big models do the heavy lifting for now.
Comment by ipsod 1 day ago
Even if you aren't coding, you really need to double check its answers. Flash 3.8 hallucinated a Keyence camera's max operating temperature for me, last week, and backed it up with "references".
It's still my favorite model for most non-coding stuff, though.
Comment by svachalek 1 day ago
Comment by esafak 1 day ago
Comment by johnsimer 1 day ago
Comment by rayiner 1 day ago
I don't know if it's the plain english or what, but I really like Grok for legal research (as opposed to code). It's got a noticeable edge in getting to the point compared to Opus 5.
Comment by dumberquestions 1 day ago
Comment by user43928 1 day ago
How representative that is of real world usage, I don't know.
In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.
Comment by attentive 1 day ago
And like that grok4.7 cache reads are more expensive than sol's (at $0.40/mil).
Comment by giancarlostoro 1 day ago
I do wonder why a frontier model does this to be honest. It still does good coding wise, but it seems strange to me. r/Claude is full of "load bearing" jokes in every thread.
Comment by Waterluvian 1 day ago
Comment by algoth1 1 day ago
Comment by fakwandi_priv 17 hours ago
Googling it returns no matches but I think it was supposed to be “live viewer”?
Comment by r_lee 13 hours ago
Comment by shawabawa3 1 day ago
Comment by tk90 1 day ago
Wonder if we'd benefit from a much more specialized + task-specific benchmarks to paint a clearer picture like this. A benchmark solely for frontend, ruby, hardware, etc.
Comment by dmix 1 day ago
Comment by ndesaulniers 19 hours ago
Lol, probably because Tesla's software stack is buildroot based. I'll bet that was in the training data.
Comment by laurels-marts 1 day ago
So yea, I find fable 5.1 writing to be excellent everywhere. I still use Sol daily though, but for things like config, quick research, fixes, code review etc. Feature work and writing is for fable 5.1.
Comment by bushbaba 20 hours ago
Comment by atniomn 1 day ago
Comment by moojacob 1 day ago
Fable 5.1 is not there quite there yet.
They need to get that Sonnet 3.5 magic back.
Comment by rfgplk 1 day ago
// `HANDLE` is an opaque kernel handle (kernel32 validates and returns 0/FALSE
// on a non-console handle); every out-param is `&mut T` to a `#[repr(C)]` POD,
// ABI-identical to the Win32 `LP*` pointer (thin non-null). The reference type
// encodes the only pointer-validity precondition, so `safe fn` discharges the
// link-time proof. (`bun_windows_sys::kernel32` declares these with `*mut`;
// redeclared locally so the legacy-conhost cursor path below is plain calls.)
or // Progress's terminal handle is the canonical `output::File` (vtable-backed
// stderr/File from `OutputSinkVTable`). The duplicate `ProgressTerminalVTable`
// from B-0 round 1 is removed; tty/ansi/winsize route through the new
// `OutputSinkVTable` slots so `bun_core` stays T0 (no `bun_sys` dep).
from src/bun_core/Progress.rsComment by imron 1 day ago
The Claudish is dead. Long live the Claudish.
Comment by sscaryterry 1 day ago
Comment by 7734128 1 day ago
Comment by fatata123 1 day ago
Comment by sscaryterry 1 day ago
Comment by joegibbs 1 day ago
Comment by qaq 1 day ago
Comment by aditya-ramabadr 22 hours ago
Comment by xmorse 1 day ago
Comment by pietz 1 day ago
Comment by octoberfranklin 19 hours ago
And the fact that Grok is the ultimate grandmaster of parallel tool-calls, routinely kicking off four or five at once. Overlapping the latencies makes a huge difference in responsiveness.
I also like how Grok is trained to print a short one-sentence descriptions of what it's doing before each step. Like an airline pilot calling out observations for the black-box recorder to hear.
Comment by Imustaskforhelp 22 hours ago
That’s a factor of half the parameters. I would be curious to see more on the focus of smaller parameters model and pushing its frontiers
Comment by petesergeant 1 day ago
Comment by Rover222 1 day ago
Comment by quater321 1 day ago
Comment by quater321 1 day ago
Comment by mrtesthah 1 day ago
Comment by DoesntMatter22 20 hours ago
Comment by throw10920 18 hours ago
Comment by Forgeties79 1 day ago
I don't want to waste money because my calculator is cracking jokes. They don't deserve their paltry 5% marketshare or whatever it is they have currently. I'm not even getting into Musk as a person or the horrid things we've seen Grok spit out on twitter. I just don't trust his companies with my data and I have seen very little evidence that it's ever the best tool for the job. I'm sure those cases exist but I can't imagine it's worth it.
Comment by StilesCrisis 1 day ago
Comment by simonw 1 day ago
Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
UPDATE: I tried again with the xAI API directly: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.
For comparison here's a fresh run against Grok 4.6: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Comment by TomGarden 1 day ago
Comment by athrowaway3z 1 day ago
It used to be a mess in various interesting ways. Now, almost every big release can draw something perfectly functional.
So the question - without a correct answer - given the prompt "Generate an SVG of a pelican riding a bicycle":
Does the user want the least lines of code to make it functional, or the best looking version?
Comment by wolttam 1 day ago
Comment by peder 11 hours ago
Comment by Mashimo 1 day ago
Comment by paimapi 1 day ago
Comment by forgot-my-pw 1 day ago
Comment by datsci_est_2015 1 day ago
Comment by MattDamonSpace 1 day ago
Comment by jfoster 22 hours ago
Has a shadow
Better shaped beak
Leg position more realistic for bicycle riding
Better feathers
Comment by kiliancs 1 day ago
Comment by daveguy 1 day ago
Comment by simonw 1 day ago
Comment by nicolamanzini 1 day ago
Comment by iankp 15 hours ago
Comment by danappelxx 23 hours ago
Comment by mrtesthah 23 hours ago
Comment by vessenes 1 day ago
Comment by mrtesthah 23 hours ago
Comment by howunfortunate 23 hours ago
I'm curious if you feel the same about re-migration of Belgians from the Congo?
Personally I think it's fine for any country to vote to control immigration as they see fit. I think Japan is a good example of a relatively xenophobic culture that deals with this fairly and thoughtfully.
Comment by thinkcontext 21 hours ago
Can't say I've ever heard anyone implying that colonialists leaving Belgium was unjust. Colonialists is actually not the right word, more like extended occupation, only slightly better than the enslavement of the Leopold II era. The Belgian's were less than 1% of the population and all but an ancillary amount worked in exploiting the native population.
Comment by CuriousRose 22 hours ago
Comment by throw10920 18 hours ago
Comment by hardbass 13 hours ago
Comment by thinkcontext 21 hours ago
Comment by bigyabai 22 hours ago
Comment by throw10920 20 hours ago
Irrelevant. HN is not the place to randomly inject flamewars about politics. It's explicitly against both the purpose and guidelines of HN.
Seems like you need to review the guidelines again, because they're pretty clear:
> Eschew flamebait. Avoid generic tangents. Omit internet tropes.
Comment by UltraSane 18 hours ago
Basic morality is not "flamewars" or "politics".
Comment by bigyabai 20 hours ago
The fact that Elon Musk's companies take contracts from the CIA and NRO is not flamebait. It's context that informs how we evaluate future SpaceX ventures.
Comment by mrtesthah 7 hours ago
Comment by UltraSane 22 hours ago
Comment by runsWphotons 21 hours ago
Comment by mrtesthah 7 hours ago
Comment by mchusma 1 day ago
4.7 is definitely slower & more expensive. It feels kind of like they really had it burn tokens to claw up the benchmarks. But it's not super clear to me whether it's above the line or not. A part of that is that it is so slow that i haven't been making fast progress today with benchmarking it.
Overall, it it gets above my intelligence line its a good release...but you can read the tea leaves and tell the Grok team thinks this was a miss.
Comment by DustinBrett 23 hours ago
Comment by pampas 23 hours ago
Comment by andsoitis 1 day ago
For example?
Comment by mchusma 1 day ago
4.6 made more mistakes than SOL or Opus overall. Gave up a lot. And in my opinion, the rate of mistakes is kind of more important than how brilliant it is.
I think 4.7 may still be better, but I was hoping for clearly Sol/Opus level and so far it just isn't there for me.
Comment by the_sleaze_ 1 day ago
Make a galaxy model search the space, create a document, argue and defend decisions, then hand it to 2.5 to implement. 4.6 was a slower less enjoyable version of that.
4.7 is better at "I want the button to cancel the jobs, dont make any mistakes" but honestly that's not what I use it's class for.
Comment by jorl17 22 hours ago
The Django code that comes out of composer2.5, to me, was insulting. Grok definitely was a step up, especially because the fast option reaaally is fast so even if it came out a bit wrong I could just whip it into perfection.
For frontend work, it's a different story. You can still tell that composer2.5 is taking the long route, but I don't think it's as egregious as with Django.
Also, composer2.5 would routinely run commands that were really dangerous and in need of proper sandboxing. Things like creating an ./uninstall.sh script with a HOME variable on which it does rm -rf $HOME. In general, when I asked composer2.5 to do things "for me", I knew a third of the initial commands would be failures, and sometimes they could be catastrophic failures (it did actually run rm -rf $HOME on what would be an actual home folder). This just hasn't happened with Grok.
I also have a bunch of vibe-coded apps I built for myself with composer2.5 and it is extremely noticeable that they hit a "this needs to be refactored as it's crumbling unto itself" line much earlier than with Grok and proper frontier models.
Comment by ActorNightly 23 hours ago
Comment by perilunar 18 hours ago
Comment by ActorNightly 17 hours ago
Comment by peder 11 hours ago
Comment by ActorNightly 6 hours ago
Comment by pclowes 20 hours ago
This is such cope.
Comment by ActorNightly 17 hours ago
Comment by merely-unlikely 3 hours ago
Not to mention the whole launching reusable rockets thing which is pretty cool too.
[1] https://archive.is/20260619220349/https://www.nytimes.com/20...
Comment by janderson215 12 hours ago
Comment by ActorNightly 6 hours ago
If you think anything Elon doing is groundbreaking, you have no idea how the world works. Recent Space X ipo showed that the launches aren't cheaper, they are just heavily subsidized. Tesla was a piece of crap until they got their model 3, the only reason Tesla succeeded with their S model is because Elon was the edgy hype dude who managed to generate enough hype to carry them through the bullshit with the car. Self driving was supposed to be solved last year, and tiny companies like Comma AI manage to build self driving systems that are in someways better than Teslas.
I bet you think Steve Jobs was a visionary as well lol.
Comment by vachina 5 hours ago
If not for SpaceX there wouldn't be gigabit internet connectivity in the middle of the ocean.
You're entitled to your own opinion to hate the guy but some self reflection goes a long way.
Comment by ActorNightly 2 hours ago
Most manufacturers already were working on hybrids, which to this day are still suprerior to EVs. Chevy Volt, outside of being Chevy, was still one of the best cars ever made for utilitarian purpose. Nobody wanted to foot the bill to do electric conversions until this was necessary.
Tesla only opened up a market segment for high end electric cars, which I guess is cool, but far from revolutionary. The model 3 was a big success only because again, it was subsidized. Meanwhile BYD actually makes cheap affordable electric cars, and we both know why they are not sold in US.
>If not for SpaceX there wouldn't be gigabit internet connectivity in the middle of the ocean.
Plenty of companies were doing geostationary orbits with satellite connectivity. SES for one.
Any more Elon slop? You realize you are defending a dude that is literally a Nazi, right?
Comment by pclowes 7 hours ago
Neuralink: to see the impact of increased independence and autonomy of a paraplegic one day after the operation
SpaceX: to quite literally approach the final frontier. Currently launches 80-90% of all orbital mass. Starlink is saving lives constantly.
Tesla: to kickstart the EV revolution and reduce fossil fuel dependence
Boring Company: to radically decrease tunneling costs applicable to all sorts of critical urban problems from transportation to utilities etc
Comment by ActorNightly 2 hours ago
* Start their own company
* Go work for a startup where they actually get paid options, and have a say in what the company does.
* Go chill at one of the big companies getting paid good salary while coasting because the work is so easy.
What they certainly don't do is go work at a company for less pay and harder work hours, all so that they can make a literal Nazi richer.
Comment by saejox 1 day ago
xAI missed its chance, Ball is on Anthropic's court.
Comment by redox99 1 day ago
Comment by jstummbillig 1 day ago
Comment by brianwawok 1 day ago
Comment by nl 20 hours ago
Elon claimed Opus was 5T in April, and I think it's fairly likely this is accurate: https://x.com/elonmusk/status/2042123561666855235
Comment by enraged_camel 1 day ago
It's phenomenal at computer use and 3D stuff. I've been using it less and less for coding.
Comment by brink 1 day ago
Comment by haellsigh 1 day ago
Comment by adventured 14 hours ago
Comment by jquery 20 hours ago
Comment by mcintyre1994 14 hours ago
Comment by cowboylowrez 12 hours ago
Comment by sneezychl 1 day ago
Best to stick with a high end model + low effort, do a manual pass on high effort and fix the bugs you know are reachable.
Comment by epolanski 1 day ago
The two models are in completely different price tiers. Astra costs 5 times as much.
It seems like all you can judge about cars would be their maximum speed on an oval.
Comment by user43928 1 day ago
Based on Artificial Analysis Cost per Task, Astra is about 2-3x cheaper than Fable 5.1 at Medium and Low.
Consequently Astra could be cheaper than Grok 4.7, depending on the task.
Comment by 01100011 1 day ago
Comment by manmal 1 day ago
Comment by Razengan 1 day ago
Comment by pac0 1 day ago
Comment by manmal 4 hours ago
Comment by Razengan 1 hour ago
Comment by imposter 1 day ago
Comment by kvirani 23 hours ago
Is the training data more valuable ? The training process ? The harness ?
I know they are all important but where are they (all the frontier labs) really pushing to get incremental gains?
Comment by pram 22 hours ago
Comment by FergusArgyll 18 hours ago
An amateur but worth reading nontheless
Comment by YZF 18 hours ago
Comment by phazeedae 15 hours ago
Comment by artdigital 19 hours ago
My SuperGrok subscription previously easily lasted me through the week even with mild coding through Grok Build. Now when I use the app 1-2 times a day to ask some questions, I’m almost running out by the end of the 7 days. It’s terrible.
I want to keep using Grok but logically it makes no sense for me to keep paying for it on the side when my quota just doesn’t last. I have also no desire to upgrade to Plus with these terrible limits, while previously I would have eaten up a $100/mo Grok plan. Rumors say SuperGrok got heavily nerfed with the SuperGrok Plus introduction, and that sounds about right to me.
I’m sure it’s a great model and I’d love to use it. I hope they get their plans under control and only only focus on Grok Bot.
Comment by hcurtiss 2 hours ago
Comment by anderber 4 hours ago
Comment by nomilk 17 hours ago
Even worse, the mere risk of quota exhaustion mid-conversation makes me not risk starting convos with Grok, instead I'll use ChatGPT or Claude (even though Claude is inferior for non-coding tasks, and ChatGPT is inferior for all tasks).
In fairness to xAI, they're a profit-motivated company like any other, so they cannot give us tokens for free or less than it costs them. Reality is we may have to simply pay a lot more if we want that Grok goodness.
Comment by artdigital 13 hours ago
Claude voice mode if you put it on Opus is now also pretty good, but there are frequently situations where I get upset at it’s responses.
Comment by dom96 1 day ago
Comment by sejje 1 day ago
For now, I doubt anyone would notice your protest if you didn't announce it.
Comment by lirolero 1 day ago
Comment by peder 1 day ago
Comment by ulfw 19 hours ago
Comment by thefourthchime 18 hours ago
Comment by OrangeMusic 16 hours ago
Comment by hersko 8 hours ago
Totally the same.
Comment by mempko 1 day ago
Comment by mempko 1 day ago
Comment by cbeach 16 hours ago
If Elon hadn’t worked with Orange Man Bad, then the Left would still be in love with him for his massive former donations to the Democrat political machine, and his work against climate change.
The whole “he’s a nazi” accusation is banal, and people are seeing through it now. That’s why we’ve moved on.
Comment by WarmWash 9 hours ago
Like the whole pizza parlor pedo basement thing, people will death grip stupid stuff because they are so desperate to manifest the worst possible image of those unaligned with them.
The problem is that it blows up in their face and just makes them look unreliable, dumb, and lost.
Musk has done so many objectively bad things that there is no need for people to dilute their reputation on fringe theories and interpretations. Pushing the nazi thing just gives Musk ammo that his detractors are so desperate that they need freeze frames and hidden context to make him look bad.
Comment by dom96 1 hour ago
If you want him to get the benefit of the doubt about his "hand gesture" then it would help if he wasn't promoting far-right parties all over Europe.
Comment by purerandomness 13 hours ago
Comment by niek_pas 16 hours ago
Comment by xerlait 15 hours ago
Comment by qwerpy 1 day ago
Excited to try 4.7. I hope they fixed the "it's not X, it's Y" that showed up in 4.6.
Comment by Theodores 1 day ago
By now AI should know of the DRY concept. But no. Hence the keys have a rounded rectangle for the key shape and another rounded rectangle for a clip path, to prevent text overflow. There are 72 * 2 = 144 identical rectangles, when just one would suffice (in the defs), with this being cloned once for the clip path, and 72 times for the keys.
I would not expect SVGO levels of optimisation (rounding numbers, that sort of thing), however, the human, if writing out the same thing for the 72nd time, might think 'is there a better way', to get the manual out. A graphics program such as Illustrator would not do that, but AI 'should' because AI.
The above is not criticism of your work, just an observation regarding AI SVG capabilities.
Comment by qwerpy 22 hours ago
Comment by Theodores 8 hours ago
There are also interesting inheritance rules with SVG, so you could define the basic shape of a key, well, several shapes, just as rects in the defs, with no stroke or fill specified.
Then, at the group level, you can then specify stroke and fill, so there could be a group of normal keys, another group for modifiers, function keys and so on.
Then there are the keys themselves, how do you clone a shape and put different text inside each clone? There are many ways to do this but I think you are on the right track using the clip path approach, albeit using the rects in the defs.
What is interesting about SVG is that artists don't care for the file format, they just see text as shapes on a page. Then programmers don't care for SVG as that is a graphic designer/artworker thing. So SVG sits in this witch-space, with only a few brave enough to wade in and do cool stuff.
Given your application, and given the fun that could be had with SMIL/JS, you could make your SVG files interactive, so you press a key and a popover tells you more about what that key does. You can even get audio working in SVG, as well as HTML popovers (in foreignobjects, as buttons, but working, nonetheless).
'Views' is another interesting SVG feature. I have a sprite sheet that uses a lot of views, where you are projecting your SVG into some type of virtual canvas, taking a 'picture' of it, and then incorporating that in something else, maybe a CSS variable.
One 'deadly addiction' is animation. Filters are another 'deadly addiction'. Why have a static and actually useful diagram, when you can animate it, move the 'camera' and add the equivalent of 27 Photoshop layers as filters to everything?
For example, supposing you wanted to show what keys to press, with there being modifiers and a sequence, e.g. 'Hello World!'. The animation for one letter could be what triggers the animation for the next 'key press' and so it goes.
My top tip of all: reposition the origin (0,0) to where it makes sense. Many objects have symmetry, so you can define one side, clone it, scale it (-1,1) and do it all around 0,0 to then translate the results to somewhere sensible.
I have found the JetBrains IDEs to be extremely useful for SVG, the preview feature is very helpful, as are the code hints.
AI sort of knows SVG, so I have had some suggestions from Google on how to build filters. These never work, but they do get you thinking. Say you wanted to use filters to add specular highlights and animated shadows to the keys, that would be fair game for AI hints on how to do it.
Comment by jjcm 1 day ago
Designs: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
Astra's build: https://html.non.io/annui/
Grok's build: https://html.non.io/Annui-grok/
Additional prompt instructions: "Add scrolling clouds behind the statues. Dynamically light the statues based on mouse position. Use diffui to generate the normal maps/depth maps/roughness maps of the objects, and to separate out the assets on to different layers."
Overall I find these models are getting good at following image as a source of instructions, but their refinement of the output varies heavily between the models. Astra's final output feels more polished, has better visual contrast, and the animations between the pages are smoother. Grok also chose to light all of the background elements, which imo overcooks it a bit.
Still though, for the price it's a great starting point.
Comment by jjcm 1 day ago
Grok 4.7: $12.60
GPT Astra: $35.00
Comment by Kuyawa 1 day ago
Comment by jjcm 1 day ago
Comment by techazard 11 hours ago
Comment by Ruphin 22 hours ago
Comment by meerita 1 day ago
Comment by vorticalbox 1 day ago
In cursor I have switch over to grok for planning a composer for coding.
Comment by brianwawok 1 day ago
Comment by vorticalbox 1 day ago
Comment by thefourthchime 1 day ago
Comment by testfrequency 1 day ago
Comment by user43928 1 day ago
Comment by brcmthrowaway 1 day ago
Comment by microsoftedging 1 day ago
"Privacy# All these models are hosted in the US. Providers follow a zero-retention policy and do not use your data for model training, with the following exceptions:
Big Pickle: During its free period, collected data may be used to improve the model.
DeepSeek V4 Flash Free: During its free period, collected data may be used to improve the model.
MiMo-V2.5 Free: During its free period, collected data may be used to improve the model.
Laguna S 2.1 Free: During its free period, collected data may be used to improve the model.
Ling-3.0-tiny Free: During its free period, collected data may be used to improve the model.
LongCat-2.0 Free: During its free period, collected data may be used to improve the model.
North Mini Code Free: During its free period, collected data may be retained and used to improve the model. Do not submit personal or confidential data. See the provider’s Terms of Use and Privacy Policy.
Nemotron 3 Ultra Free (NVIDIA free endpoints): Trial use only — do not submit personal or confidential data. Your use is logged for security purposes and to improve NVIDIA products and services. The logged session data for improvement purposes is not linked to your identity or any persistent identifier. For more information about data processing practices, see the Privacy Policy. By interacting with this endpoint, you consent to the collection, recording, and use of such information and the NVIDIA API Trial Terms of Service."
Comment by BeetleB 1 day ago
Comment by thehamkercat 1 day ago
https://openrouter.ai/deepseek/deepseek-v4.1-flash?endpoint=...
Comment by thehamkercat 15 hours ago
Comment by riquito 1 day ago
Comment by nicce 1 day ago
Comment by thehamkercat 15 hours ago
Comment by sparkling 1 day ago
Comment by RussianCow 1 day ago
"order": ["relace", "coreweave", "novita", "baseten", "together"],
"allow_fallbacks": false
It still won't be quite as high as you'd get by just using DeepSeek because occasionally a request will fail and you'll get routed to a backup provider with nothing cached, but it's close enough not to matter in most instances.But I can't argue with the lower off-peak pricing when using DeepSeek directly. The downside is they train their models on your input, which might be a deal-breaker for many users (as it is for me).
Comment by drewnick 1 day ago
Comment by simlevesque 1 day ago
Comment by parineum 1 day ago
Comment by meerita 1 day ago
Comment by includenotfound 1 day ago
Comment by meerita 1 day ago
Well, at least I spent lots of dollars, and I had to use those models the same way I am using local and cheap models, with the same results.
Comment by includenotfound 17 hours ago
On the other hand, Grok and GPT finish these tasks in <5min with no issues, and significantly better output.
Comment by meerita 15 hours ago
Comment by _s_a_m_ 1 day ago
Comment by yipinwong 1 day ago
GLM or Kimi are better for my own personal projects. DS? uhm. it just keeps doing dumb crap
Comment by joegibbs 20 hours ago
I told it to compose an image (putting headgear on top of a head) - kept getting it completely wrong, generating new headgear, getting that wrong and screwing up the scaling.
I told it to diagnose a webhook issue that was happening in production from a local environment and it kept giving me moronic answers like that environment variables weren't set (despite me telling it that the values WERE set in production).
Comment by DoesntMatter22 20 hours ago
Comment by dozerly 20 hours ago
Comment by csomar 19 hours ago
I've tried the latest models across OpenAI, Claude, Chinese, etc. They just do stuff. That's not how work is though. You want them to do specific work, at which point it's a real hassle to follow up on all the garbage they have been outputting.
/long rant
Comment by sarjann 1 day ago
Unless they produce the same token output on the face of it, it looks like they're trying to cover for 4.7 not having good model perf?
Comment by trentor 1 day ago
Comment by notduckrabbit 1 day ago
Comment by sourcecodeplz 1 day ago
- grok 4.6 (xhigh): 97M (for 44 score)
- grok 4.7 (xhigh): 240M (for 46 score)
Comment by everfrustrated 1 day ago
Comment by notduckrabbit 1 day ago
Comment by WarmWash 1 day ago
Comment by johnfahey 1 day ago
Comment by stiltzkin 1 day ago
Comment by GodelNumbering 1 day ago
From their headline comparison:
Grok: $2/$6 per million
Fable: $10/$50 per million
What this doesn't say: Grok costs 0.50/M cache read, Fable $0.25/M cache read
Long running agentic workflows are dominated by cache reads.Just makes Grok sound deceptive, and more importantly, reliant on user's lack of understanding of costs aka predatory (which in turn is more infuriating)
Comment by sourcecodeplz 1 day ago
Comment by bastawhiz 1 day ago
Comment by shdtabasum 1 day ago
Comment by xquce 1 day ago
Comment by MoreThanMe 3 hours ago
Comment by pampas 23 hours ago
Grok 4.7 is near the top of the board. A significant improvement over Grok 4.6 but still not as good as Gemini 3.8 Flash which is very cheap and fast too.
Comment by ls1911 1 day ago
Comment by wg0 14 hours ago
Comment by sbseitz 22 hours ago
Comment by thefourthchime 18 hours ago
Comment by totallymike 10 hours ago
Comment by gslepak 1 day ago
Comment by daquisu 1 day ago
It is the same multiplier for Sol with subscription. For Astra though the multiplier is ≈20x, so half of Sol usage.
For Claude it seems to be ≈40x too for Opus, but less for Fable (similar to Astra in GPT).
All on the most expensive plan. Previously, Grok usage escalated linearly from the $100 plan to $300 plan. That would be a really good $100 plan if it is still true.
Some sources:
1. https://x.com/kunchenguid/status/2098256018836963382
2. https://x.com/stevenzhang/status/2092110386569089311
3. https://github.com/openai/codex/issues/43731
Comment by nwienert 1 day ago
It reset just a few hours ago and I've been running it, couldn't be more than 15 sessions none more than an hour long:
---
Session usage: no model calls yet in this session.
Weekly limit: 46%
Next reset: September 27, 23:20
---
I actually have to believe my account is messed up tbh, it's so bad. For reference I've ran 12 fable and some ~40 Opus sessions since reset yesterday on a CC account, at least 5x more usage by my estimate:
Current week (all models)
28% used
Resets Sep 28 at 5am (Pacific/Honolulu)
---Ok looking at it more, Grok and Grok Build just really suck. They are about 10x less token efficient, often using 200+ tool calls in a row for what are not even big tasks where Opus would use 5-10. Their cache hit rate is worse, and two sessions got into basically unnecessary loops costing a solid quarter of the entire week. And this was on smaller tasks as I tend to use it for easier things.
Comment by daquisu 23 hours ago
I can't help much more than that, I did that research in the last few days, but I never used Grok myself.
I pay for GPT, Claude and Gemini. Last week I consumed all my quota on two of them, so I wondered which next subscription I would pay for if needed.
Comment by nwienert 21 hours ago
Comment by everfrustrated 1 day ago
For me and what I’m doing that’s insanely good value.
I find grok build chews through my SuperGrok sub very quick - but I think that is due to it having the 500k context window which uses more credits. Cursor limits it to 256K (tho I see in today’s update for Grok 4.7 there’s now a toggle for context size).
Comment by artdigital 19 hours ago
Normal SuperGrok barely lasts me through the week with very mild usage and no coding. The sentiment around SuperGrok Plus is also not great and I haven’t seen someone saying they’re happy with it yet.
SuperGrok Heavy is $300/mo, so you could get a full ChatGPT Pro and Claude Max 5x for that price. That’s so far out of my budget for a single provider I haven’t bothered trying it.
I still have SuperGrok through X Premium+ but will downgrade that next billing cycle
Comment by andreyvit 1 day ago
Comment by thefourthchime 1 day ago
Comment by nwienert 1 day ago
Comment by swalsh 1 day ago
Comment by becquerel 1 day ago
Comment by sparkling 1 day ago
Astra for deep dive investigations, Sol 5.6 at mid-level for day to day tasks, Grok 4.6 via Cursor for routine and low complexity tasks.
Comment by maz1b 1 day ago
Comment by avazhi 1 day ago
There for awhile it seemed like we’d have 3 big competitors but then Grok 4.2 or 4.4 was just diabolical while OAI and Claude continued their significant improvements. Grok was/is so bad that I was convinced musk was gonna shut it down and just fund Anthropic compute once they reached their compute agreement.
Comment by 6thbit 1 day ago
Comment by asdfsa32 22 hours ago
But honestly, it is because numbers are like people; torture them enough and they'll tell you anything.
Comment by sourcecodeplz 1 day ago
Output tokens from Intelligence Index:
- grok 4.6 (xhigh): 97M (for 44 score)
- grok 4.7 (xhigh): 240M (for 46 score)
Comment by oh_no 1 day ago
Comment by simonw 1 day ago
Comment by AM1010101 1 day ago
Comment by ssutch3 1 day ago
Comment by forgot-my-pw 1 day ago
Comment by ssutch3 1 day ago
Comment by everfrustrated 1 day ago
Comment by alansaber 1 day ago
Comment by epsteingpt 21 hours ago
Comment by claaams 21 hours ago
Comment by oh_no 1 day ago
Comment by c0rruptbytes 1 day ago
Comment by mh- 1 day ago
Comment by sidgtm 1 day ago
Comment by aschobel 1 day ago
Comment by guywithahat 1 day ago
I'm excited for 4.7 although I share skepticism with other users whether 4.7 will be significantly better, since they didn't raise the price.
Comment by XCSme 23 hours ago
https://aibenchy.com/model/x-ai-grok-4-7-medium/#showcase=dd...
Comment by XCSme 23 hours ago
https://aibenchy.com/model/x-ai-grok-4-7-xhigh/#showcase=2f9...
Comment by jascha_eng 1 day ago
Comment by karma_daemon 22 hours ago
Comment by MuffinFlavored 1 day ago
Is there a metric for like... time taken when comparing these two? I see score and cost.
If Fable5.1 can knock it out more quickly on low but Grok4.7 might take twice as long to stumble through a problem (and leave behind a bunch of yucky comments or un-needed extra unit tests), are they really comparable?
Or like... the "quality" of the solution? "It works" versus "it's unmaintainable/very messy/hacky".
Comment by inshard 1 day ago
Comment by mpalczewski 1 day ago
Comment by gaigalas 1 day ago
Comment by brcmthrowaway 1 day ago
Comment by thih9 1 day ago
But also Xai doesn’t seem to care about user experience and long term support.
Comment by eknkc 1 day ago
For daily one off questions I prefer it because it is fast enough and I like the way it responds. I also use it for basic research like “find me a battery drill for this and that”.
Kimi and GLM feel extremely coding oriented. I use them for code reviews basically. I hate the way Anthropic models talk. GPT takes too much time and effort for that kind of stuff for some reason.
Grok happened to be a nice middle ground.
Comment by thefourthchime 18 hours ago
Comment by jpadkins 22 hours ago
Comment by nimchimpsky 22 hours ago
Comment by brandonagr2 1 day ago
Comment by venzaspa 1 day ago
Comment by swozey 1 day ago
As a technical point of reference to compare against other llm stuff, sure, I'll glance at a report or benchmark but I really couldn't care less about anything to do with the project and it could blow other options away and I wouldn't touch it.
Comment by ElectronCharge 1 day ago
You probably shouldn't cut off your nose to spite your face.
Comment by totallymike 22 hours ago
Comment by mempko 1 day ago
What's superficial about refusing to use a product from someone like that? Or are you one of those 'technology isn't about politics' people? That's a superficial take if you ask me.
All technology is political, and understanding that is a deep, not superficial take. It requires systems thinking which unfortunately many people building technology seem to lack, despite software being a sophisticated complex system.
Comment by Razengan 1 day ago
Comment by Klathmon 1 day ago
I'll never use an xAI product.
Comment by jfoster 22 hours ago
Comment by Klathmon 21 hours ago
I'm still never going to use an xAI product.
Comment by Razengan 21 hours ago
If a shitty murderous tyrant builds some roads the citizens can still use those roads while protesting against the tyrant
Comment by Hikikomori 15 hours ago
Comment by nimchimpsky 22 hours ago
Comment by Saline9515 1 day ago
It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.
Comment by xmorse 1 day ago
Comment by Saline9515 1 day ago
Comment by samtheprogram 1 day ago
Comment by marwatk 1 day ago
- it allows different models within one session via roles (I only have API, so pay per token)
- it's much more likely (ime) to use the LSP over grep for determining how code fits together
But I agree a 20k+ starting context is way overkill.
I find it's very hard to get information on harnesses people are using. I have to stay model agnostic so I avoid claude, codex, cursor, etc. I've used and tried opencode, which worked well, but obviously lacks the above features.
Does anyone have a resource for following what people are actually being productive with? With so much vibe going on it's hard to separate the wheat from the chaff.
Comment by xmorse 1 day ago
Comment by raincole 1 day ago
Comment by Saline9515 1 day ago
Comment by polytely 1 day ago
Comment by raincole 1 day ago
Comment by unrvl22 1 day ago
Comment by kristofferR 1 day ago
Comment by Jcampuzano2 1 day ago
This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.
Comment by user43928 1 day ago
That said, I don't expect them to benchmark Astra in their Cursor harness given the situation.
Comment by Jcampuzano2 1 day ago
Comment by oh_no 1 day ago
Comment by kristofferR 1 day ago
If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.
Comment by andsoitis 1 day ago
Comment by kristofferR 1 day ago
Comment by Jcampuzano2 1 day ago
Comment by mh- 1 day ago
Comment by scottyah 1 day ago
Comment by kristofferR 1 day ago
Comment by scottyah 23 hours ago
Comment by Iolaum 1 day ago
Comment by Jcampuzano2 1 day ago
Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.
Comment by babelfish 1 day ago
Comment by Jcampuzano2 1 day ago
Comment by ryeguy 1 day ago
Comment by bluecalm 1 day ago
https://x.com/elonmusk/status/2102082011233931762?s=20
so it's likely about usage in Cursor specifically.
Comment by DavCreator 1 day ago
Comment by babelfish 1 day ago
Comment by szamski 15 hours ago
Comment by usumgallu 1 day ago
Comment by Tsarp 1 day ago
Comment by forgot-my-pw 1 day ago
Comment by rvz 1 day ago
Comment by jcims 1 day ago
Comment by kridsdale3 1 day ago
Comment by user43928 1 day ago
Comment by TylerE 1 day ago
Comment by lumirth 1 day ago
Comment by user43928 1 day ago
That it isn't the most efficient way to achieve the same end result is irrelevant.
Comment by bluepeter 1 day ago
Comment by nicolamanzini 1 day ago
Comment by mempko 1 day ago
Comment by blactuary 1 day ago
Comment by 13415 1 day ago
Comment by Shiggy_ 1 day ago
Comment by DaSHacka 1 day ago
Comment by felixgallo 1 day ago
Comment by myko 1 day ago
Comment by TylerJaacks 1 day ago
Comment by mavamaarten 1 day ago
Comment by totallymike 18 hours ago
Comment by nostrebored 17 hours ago
Many people are allergic to “AI safety” as they perceive it as an attempt to deceive them. When they ask a factual question and get back an unfactual answer, it makes them upset. None of the people I know who feel this way are searching out CSAM. They accurately perceive that there is a team that wants answers to come out a certain way.
For instance, I’m loosely connected to people who care about animal safety in AI research. There are absolutely people spending their time pushing a narrative about _the_ ethical way to interact with animals. Feeling this not being forced onto you is kind of nice.
Comment by totallymike 10 hours ago
Comment by t1E9mE7JTRjf 18 hours ago
Comment by totallymike 17 hours ago
Comment by johnnyApplePRNG 1 day ago
Comment by eleventen 1 day ago
Comment by big_toast 1 day ago
These comments don't stay up much anymore and I can't tell if it's structural to the forum (flag weight + statistical mechanics of votes + guidelines) or if it's the userbase sentiment.
Comment by oulipo 1 day ago
Comment by eleventen 1 day ago
But I think it represents real malaise in the community. It's not a moderator plot, people here really just don't care and might even support this.
We really are in the minority of opinion for giving a damn about liberal democracy.
Comment by big_toast 1 day ago
Between the guidelines + user thoughts (e.g. repetition, low novelty/new info), there's other reasons these types of replies might end up dead.
I am worried that it leads to people self selecting to other forums biasing the remaining userbase vote/vouch/flag distributions. In an exit vs voice situation, the voice kinda dies out. Then we end up other-izing people and homogenizing our communities.
But I concede it's also possible that the minority opinion issue could be the core driving force.
Comment by hdhdjdif 1 day ago
Comment by jesse_dot_id 1 day ago
Comment by drop_star 1 day ago
Comment by ForrestN 1 day ago
Comment by moolcool 1 day ago
Comment by vb-8448 1 day ago
Comment by eleventen 1 day ago
Nobody both worked and spent their money to get Trump elected like Musk. 300 million to his 2024 campaign [1]. DOGE. On-stage endorsements. Nobody even came close.
No, other big labs are not "innocent little virgins", but they're not even in the same solar system of harm as Musk. To hand-wave at the differences is to permit them.
[1] https://www.opensecrets.org/2024-presidential-race/donald-tr...
Comment by unsupp0rted 1 day ago
Comment by mlindner 1 day ago
Comment by TheOtherHobbes 1 day ago
Handing corporate code secrets to his AI model is... unusually trusting.
Comment by mlindner 1 day ago
And methane is a large percentage of all power production in the US. So again that also applies to all the other data centers. (And FWIW they've been winding down and shutting down the on site methane generators.)
And no corporate code was handed to AI models.
Comment by moomin 1 day ago
Comment by grokgrokgrok 1 day ago
Comment by toader 1 day ago
Comment by ctrlkctrls 1 day ago
Comment by toader 1 day ago
Comment by chris_money202 1 day ago
Comment by zamalek 1 day ago
Comment by thereitgoes456 1 day ago
Even with his successes (Tesla, SpaceX) he has built them up in large part by bending levers of government to his advantage.
Comment by sssilver 1 day ago
Can you provide specific examples of where Elon has bent the levers of government?
Comment by voidfunc 1 day ago
So what? Thats called being a maverick. He is very very good at executing on making money which is the point of business.
Comment by andsoitis 1 day ago
Also pushing technology forward.
Comment by redox99 1 day ago
Comment by blisterpeanuts 1 day ago
Comment by redox99 1 day ago
Anthropic: 1.25B/month
Google: 0.92B/month
Unnamed customer starting in december: 1.1B/month
Starlink monthly revenue is ~1.5B/month
Comment by chris_money202 1 day ago
Comment by redox99 1 day ago
Comment by chris_money202 1 day ago
Comment by brandonagr2 1 day ago
Comment by nozzlegear 1 day ago
Comment by keeda 22 hours ago
Comment by oulipo 1 day ago
Comment by ls612 1 day ago
Comment by Romanulus 1 day ago
Comment by jackfischer 1 day ago
Comment by nibbleyou 1 day ago
Comment by estearum 1 day ago
If anything, they voted for reduced debt burden and they got the opposite. DOGE failed at pretty much every single one of the goals that the public arguably gave it a mandate for.
Comment by serbuvlad 1 day ago
Ah, yes, democracy!, except for when the public is wrong.
Who decides when the public is wrong? We do! Who decides "what the public voted for"? We do! So we are the rulers? No, of course, not, this is democracy.
You want to become the decider of when the public is wrong and of what the public voted for? TYRANT! TYRANT!
Comment by estearum 1 day ago
This is simply epistemologically incorrect. It's obviously incorrect in this case because voters writ large do not have any idea how the government is administered and how to improve it, so even if they claimed to be voting for that, it would not necessarily be an endorsement of any particular approach.
More specifically we know it's not true in this case because there are polls. Voters didn't even claim to care about this! "How the government is administered" was not a high salience issue to voters. Simple as that.
Nonetheless, I didn't suggest anything about overriding their votes. It sounds like you have some sensitive spots to work through (someone obliquely criticizing your idol for sucking at his job?)
Comment by verdverm 1 day ago
half of voters don't pay any attention to politics until the week or two before voting
Comment by maelito 1 day ago
Comment by fourseventy 1 day ago
Comment by KyleTheDev 1 day ago
Sheep often like to think themselves the wolf or coyote, it would seem.
Comment by jml78 1 day ago
Fuck, it is like the denial around Jan 6th. Those idiots we’re live streaming that shit. I watched it go down live. Now they say they weren’t violent.
We can’t have discourse when we have legit video evidence and people refuse to open their eyes and choose to deny reality
Comment by sejje 1 day ago
Which Nazi ideologies do you think he embraces? How do you reconcile all the Nazi ideologies he rejects?
Comment by nancyminusone 1 day ago
Comment by butlike 1 day ago
"You're on, bro"
Comment by oulipo 1 day ago
Comment by AtlanticThird 1 day ago
Comment by thoman23 1 day ago
Comment by jmward01 1 day ago
Comment by andsoitis 1 day ago
Comment by jmward01 1 day ago
Comment by solid_fuel 1 day ago
Comment by sejje 1 day ago
Comment by simianwords 1 day ago
The personality is bland and it doesn’t work nearly as hard or even tries to help.
Comment by Capricorn2481 1 day ago
I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.
Comment by artemonster 1 day ago
Comment by xutopia 1 day ago
Comment by artemonster 1 day ago
Comment by Paracompact 1 day ago
Comment by xutopia 1 day ago
There are ample reasons to believe that Elon Musk is running his mother's account and that the photos weren't even real putting in question that he even had a birthday party.
If you ask Grok about what this means it will always take Elon's defence. It will vehemently deny that Elon would be capable or willing participant of such a thing even if you point out that he faked being a world class gamer, buying accounts that had done all the work and showing none of the skills when live-streaming.
Comment by jpadkins 22 hours ago
Yeah, I'm sure the guy has time to run his mom's social media account. What's your reasons or evidence? Did you consider maybe it's a social media manager one of them hired running the account?
Comment by ethagnawl 1 day ago
Until you ask it to start generating horrific imagery and then it's best in class.
Comment by Shiggy_ 1 day ago
The value of the internet is that people can share whatever they want, and use software how they want. This will mean that some people will abuse that. This is the tradeoff of a free society.
Comment by raincole 1 day ago
Sounds like a plus. Guess I will give Grok another try...
Comment by slowin 1 day ago
Comment by finnjohnsen2 1 day ago
Maybe I'm in some kind of bouble but I have never met or talked to anyone who has used Grok.
Comment by Brendinooo 1 day ago
Not sure if I'll hold the subscription but I could see myself working with it more.
Comment by chronogram 1 day ago
Comment by leftbehind 18 hours ago
Grok 4.6/4.7 feel like you are less tokens because Cursor gives you double the quota subsidising Grok specifically - there is a separate Grok/Cursor model quota bar in addition to non . Its actual token drain if you are paying API rates is generally far higher than Fable, Astra et al. You are not using less tokens.
It is weird in some reward way, it will frequently, at least for me, complete the task in the laziest way possible,Technically it is done, but that's about it.
For example a user registration system ended up with users being able to log in as anyone because the authn was a cookie set with the user's pkey id, unsigned. Nearly anything it outputs technically works but if you give it a GLM or DS pass you will find dozens to hundreds of vulnerabilities of varying hilarity (Fable refused, Opus refused).
I have seen it write SQLI-vulnerable code, asked it to review a file without saying what's wrong and it did not catch it (fresh session, single file @-tagging in harness). Again, technically, the code works and does what the prompt asked, its just handing in some of the laziest copied homework I have ever seen.
If you specifically call out to use prepared statements it will typically put a plaster over this however I have also seen Grok 4.6 use prepared statements by concenating the user input raw into a statement then executing it with no ? or named replacements, rendering it somehow a SQLI-vulnerable prepared statement. One med it added a OR 1=?, replaced the ? with another 1, then said OK you are using prepared statements now.
Comment by cvwright 1 day ago
As a chatbot it’s totally fine, virtually indistinguishable from Gemini or ChatGPT or Claude.
For coding it’s… okay. I tried 4.6 and it feels similar to Opus from 12 months ago, or maybe Sonnet from 9 months ago. YMMV.
Comment by moomoo11 1 day ago
Comment by dofm 1 day ago
It's just an observation but so far a pretty solid correlation. Musk has so severely poisoned the well in terms of his UK reputation that the only people who are open about using Grok are... well, wankers is as good a word as any.
FWIW among the AI-using people, it mostly goes Claude Code, then Codex, then whatever runs on their Mac. The only Cursor user I knew has jumped ship to OpenCode.
Comment by soaaa 1 day ago
Comment by Invictus0 1 day ago
Comment by jpadkins 22 hours ago
Comment by Invictus0 20 hours ago
Comment by DonHopkins 1 day ago
Comment by enraged_camel 1 day ago
Comment by outside1234 1 day ago
Comment by BoumTAC 1 day ago
Comment by nostrebored 1 day ago
Comment by BoumTAC 1 day ago
I like to follow them and look for benchmark for each LLM release.
Comment by nostrebored 1 day ago
Comment by zug_zug 1 day ago
I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?
Comment by sejje 1 day ago
I think it's winning on UI for normies (grok bot) and they made some claims about being pareto SOTA (lowest cost per task completed) a while back with 4.6.
I find it to be a perfectly capable model for implementation (there are many in this class--deepseek flash, spark1.3, luna, etc). I find the usage to be very generous w/ supergrok. I find the model to be just fine for 90% of what I want to do, but I use a smarter model to plan complicated things.
Comment by zug_zug 18 hours ago
I'm judging on benchmarks, and whether anybody or any company I know has ever suggested using it (not yet).
Comment by sejje 18 hours ago
I don't personally make my judgements based on how many other people mention a thing, but if that gets your code written, by all means.
Comment by swalsh 1 day ago
Comment by grim_io 1 day ago
Comment by puszczyk 1 day ago
The voice is the same AI slop as the others imho.
(This is about Grok 4.6, I didn't test 4.7 yet).
edit: clarified I mean agentic coding tasks
Comment by svachalek 1 day ago
Comment by puszczyk 1 day ago
Comment by Shekelphile 22 hours ago
Deepswe results show that grok 4.6 is more expensive per-task and consistently scores worse than: luna xhigh, glm 5.3, astra low, sol high/xhigh, opus 5 medium.
Grok also used almost 3x as many tokens/turns to complete tasks than all of those models (besides luna), so it takes way more time to complete a task.
There isn't much reason to use Grok at all, it's gotten better but it's still worse than every other player in the field, which shouldn't be a surprise considering until about a year ago they were just buying tokens from other providers and pretending it was their own model.
With gpt-6 luna and sol coming tomorrow it's going to look even worse too, especially if new luna retains the same dirt cheap pricing that 5.6 luna has.