Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users
Posted by tedsanders 12 hours ago
Comments
Comment by heaney-555 9 hours ago
Comment by vitorgrs 1 minute ago
Comment by daemonologist 8 hours ago
Comment by johnsmith1840 8 hours ago
Comment by ramraj07 4 hours ago
Comment by johnsmith1840 3 hours ago
Comment by ekidd 1 hour ago
There's no reason why Google's public stuff is this stale, overpriced and underwhelming. But at least until their next round of models drops, even calling them a "frontier lab" is starting to feel like a stretch. Which is weird!
Comment by eru 35 minutes ago
Especially considering that in the 2010s they were The Big AI company, especially after buying out boutique shops like DeepMind.
Comment by willy_k 43 minutes ago
Comment by tokioyoyo 7 hours ago
Comment by stymaar 7 hours ago
Comment by theptip 26 minutes ago
Email is a good window but lots of people talk in more depth with a chat bot.
Search is a good “purchase intent moment” but people are telling a chat bot exactly what they care about in a purchase, rather than making indirect search terms and reading review sites.
Comment by tokioyoyo 7 hours ago
Comment by livinglist 2 hours ago
Comment by keeda 3 hours ago
To expand: it seems inevitable that Google's SERP format will be replaced with a conversational / chatbot / agentic interface, which equalizes the playing field for all chatbot providers.
This is because you can stuff in much fewer ads into a chat interface compared to SERPs. (They could try stuffing more ads but that would likely just push users more to the competition who have a much lower baseline on which to show growth.) As such, Google would be forced to progressively nullify its own invincible firehose of ad revenue as they deprecate SERPs in favor of AI overviews.
Comment by red_green_yell 42 minutes ago
There's no way OAI has a long term advantage over google in replacing the search engine experience.
Google's one major weakness is it will face the innovator's dilemma as their core search revenue gets cannibalized. But they seem to have been able to get their entire org to recognize that AI is an existential threat so at least that's a good sign.
Comment by eru 33 minutes ago
Comment by willy_k 41 minutes ago
You’re painting a dichotomy that doesn’t exist and hasn’t for a good while.
Edit: I missed your final paragraph. I still disagree with this argument though, people still want to go to websites.
Comment by famouswaffles 6 hours ago
The median LLM query isn't significantly costlier than web search.
Comment by virgildotcodes 6 hours ago
Comment by mikeshi42 5 hours ago
Comment by kristofferR 6 hours ago
Comment by johnsmith1840 3 hours ago
Google is going to be able to push cost down further and for longer than anyone else. They also rightfully believe this is a company ending gamble if they fail they could be a second tier player for decades.
So better inference margins, far better operation experience serving cheap AI at massive scale, a massive warchest, and the fear of complete company irrelevence.
That's going to be one hell of a company to beat for free tier LLMs. Also chatgpt is impressive but it's notthing like google search quite yet. On top of that AI labs must use google's product for their AI.
They also own more data by an exponential margin. Anything but dominance of free AI on google's side would mean they are just so incompetent they deserve to fail.
Comment by Tostino 8 hours ago
Comment by fragmede 3 hours ago
Comment by gavinray 8 hours ago
Because everyone now outsources much of their thinking and researching to LLM's, our collective culture + brain is shaped in a cyclical manner by using them.
It's the mechanical homogenization of culture and groupthink.
Comment by crab_galaxy 3 hours ago
TBH I don’t find it useful at all for personal use. It’s totally soulless for creative ventures and absolute dogshit at researching the things I want it to be good at (I.e. planning a vacation or finding new music).
All that combined with the social stigma makes me feel pretty skeptical that it’s some kind of pop culture shaping mechanism, at least not for a few more years.
Comment by jimbob45 32 minutes ago
Comment by fragmede 2 hours ago
Comment by in-silico 6 hours ago
That would explain a lot of the terrible AI/LLM takes online.
Comment by heaney-555 6 hours ago
Comment by qingcharles 2 hours ago
Comment by fragmede 3 hours ago
Comment by thorum 7 hours ago
They took away the button a few months ago and are now putting it back.
Comment by heaney-555 7 hours ago
Perhaps I should have said "proper access".
Comment by aryehof 46 minutes ago
It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?
Comment by maxipoo 21 minutes ago
Comment by firasd 9 hours ago
So 5.6 Luna is just their next version of what they used to call 5.5 instant tier
And 5.x instant models were never much to write home about anyway so the default ChatGPT free model hasn’t been particularly distinctive since 4o
Comment by freakynit 1 hour ago
Comment by ElijahLynn 9 hours ago
Comment by miki123211 5 hours ago
I feel that work is basically split into two tiers, hard (which requires a good model and lots of reasoning by definition) and relatively easy (which won't consume much of my limits despite a great model and reasoning, so I may just as well keep it on Sol High).
Comment by majormajor 3 hours ago
This equation significantly changes if you're paying API prices vs on a subscription. (Such as if you're integrating it into a different product, where it becomes worth it to figure out what the cheapest you can leverage is.)
Comment by redox99 4 hours ago
Luna: Can be useful if price sensitive, always use with AT LEAST high effort. But codex plans are very generous, so just ignore it.
Terra: Forget it exists
Sol: Just use this. Medium can work for specific edits. Otherwise just use high or xhigh.
OpenAI basically agrees with this, and the slider gives you those options.
TL;DR: Just use sol medium/high/xhigh
Comment by drivebyhooting 4 hours ago
Comment by redox99 4 hours ago
Ultra: Might be useful, the one time I tried it, it was extremely wasteful in terms of tokens. Definitely not a daily driver.
There's also max effort, I think it can be useful but the gains compared to xhigh are quite small.
I think max and ultra are only worth it when the others fail.
Comment by drivebyhooting 4 hours ago
I don’t know if this is a good workflow.
Comment by nojs 6 hours ago
Comment by skybrian 9 hours ago
It seems like giving it a time limit or a budget in dollars would be clearer, though?
Or, keep searching until I come back to the computer and ask about progress.
Comment by pllbnk 9 hours ago
Comment by minimaxir 9 hours ago
Comment by vanuatu 9 hours ago
Comment by stymaar 7 hours ago
Comment by vanuatu 4 hours ago
very nontrivial problem knowing when to stop and making assumptions
Comment by dbbk 6 hours ago
Comment by oceanplexian 5 hours ago
Comment by redox99 9 hours ago
Comment by Jtarii 9 hours ago
Comment by redox99 8 hours ago
Comment by Marha01 39 minutes ago
Comment by 2sk21 8 hours ago
Comment by Sammi 8 hours ago
Comment by awakeasleep 9 hours ago
Comment by timpera 9 hours ago
Comment by egorfine 6 hours ago
Sadly it's not available anymore on the updated desktop app (formerly Codex) and the previous desktop app (formerly ChatGPT) is abandoned.
Comment by customguy 4 hours ago
Comment by ilaksh 9 hours ago
This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
Which I think is a fair interpretation of the term. They are general purpose intelligence in that you can get help from them about almost anything. They are not like narrow single purpose AI models.
I don't think we need that term to mean "can completely emulate a human" or "can do every task any human on earth can do as well as them".
It also needs to be differentiated from ASI with godlike powers many times greater than human.
Comment by orbital-decay 1 hour ago
>artificial general intelligence (AGI)—by which we mean highly autonomous systems that outperform humans at most economically valuable work
Comment by kkoncevicius 9 hours ago
Comment by stymaar 7 hours ago
Comment by klibertp 6 hours ago
It is a huge leap. Now we can start talking about "intelligence" at all - we really couldn't before. That we're still hovering barely above the starting line is a separate matter (also worth noting, of course).
Comment by stavros 4 hours ago
Comment by ilaksh 4 hours ago
Comment by stavros 3 hours ago
Comment by hluska 3 hours ago
Comment by kubb 9 hours ago
Well, the models are smart enough to point out why this is wrong.
Comment by lostmsu 2 hours ago
Comment by andai 8 hours ago
Wasn't there a thread the other day about how very few humans on earth can understand the new math proofs?
Although "with sufficient study" vs "not even with unlimited study" are probably worth distinguishing there.
Comment by ilaksh 7 hours ago
Comment by taytus 7 hours ago
I cannot believe the comments I read on HN nowadays. ZERO critical thinking.
Comment by lukevp 4 hours ago
Comment by 1saadcodes 1 hour ago
Comment by freakynit 1 hour ago
The value is shifting up the stack... make core intelligence free, then monetize the ecosystem built on top of it .. kinda similar to how the internet itself is free, but platforms and apps capture the value.
I think we'll see a huge push towards connectors for work and personal tools, along with much deeper os level integration. That's where the long-term moat is, not the base model itself.
Comment by judge2020 1 hour ago
Comment by colingauvin 10 hours ago
I expect a few things to happen in the next year:
1) Exclusive MCP server deals/API integrations
2) Significant switch to B2B marketing, even moreso than we've seen before, with API interfaces being paid and chat-client interfaces becoming more and more free, perhaps just with limits more on integrations or data visualization/analysis
3) US restrictions on B2B contracts with non-US hosted models that do any sort of contracting with the government
Obviously there's a bunch of stuff I'm not foreseeing. But it really does feel like the bottom of the market is collapsing into free. I assume OpenAI and Anthropic think their next generation of models will restore their halo tier status and that the cash burn is justified to just get there, but this has to really mess up IPO plans.
Comment by redox99 9 hours ago
1) Back then, even as a free user you'd be able to use the strongest model (even if with tight limits). Now, you need to pay to use Sol, and you need to pay to use Opus or Fable. It does seem fairly premium in that sense. Idk about 5.6 Luna, but the previous Instant was really bad, even for very casual users. It would hallucinate non stop.
2) When $100 and $200 per month plans launched, they were received as outrageous even here. Nowadays they are pretty common among power users.
Comment by user43928 8 hours ago
Back when coding for me still meant copy-paste from the web version, it was only worth the $20/month for me.
They only added the $100 Pro plan in April during GPT 5.4 times.
Today I happily pay $400/month for Codex and Claude Code.
Comment by thejazzman 5 hours ago
Comment by colingauvin 9 hours ago
This is kind of what I'm saying though. Bottom has fallen out, differentiation is just can you be much more premium than the competition. Currently that remains unanswered.
EDIT: I'm basing this off the assumption that for chat, premium is not a point of differentiation at all. For coding/analysis, it is.
Comment by davidguetta 8 hours ago
Comment by kingstnap 9 hours ago
Maybe Luna efficiency gain was actually significant enough that putting all the free users and giving them super generous limits makes sense.
They might be doing this to improve the messaging of AI among causal users since right now there is a huge amount of datacenter backlash in the US due to AI grievances.
Maybe they have too much excess capacity or they really want to juice token numbers and market share on their dashboards for marketing.
I also wonder if being given access to an actually a decent model like luna with actual thinking budget instead of brainless "instant" modes will start to make causal users understand the real capabilities of these models.
Comment by planb 8 hours ago
Comment by deanc 9 hours ago
Comment by ToValueFunfetti 9 hours ago
Comment by mkozlows 9 hours ago
Comment by kingstnap 9 hours ago
All of this is of pretty minor importance though. You can't read as many tokens as a subcription can produce so more chat is not the value add nor super important.
I mean there are literally so many providers for free chat if you are willing to use several seperate apps.
The real value in these subs is using codex cli, much like the real point of anthropic subs is using claude code. Because agentic work actually does require a lot of tokens.
Comment by simianwords 9 hours ago
Definitely this. The recent 80% discount was a reaction to Deepseek's update so that they still position near the frontier. My theory: Luna has always had a much higher efficiency. You do know that the model didn't get faster after the discount?
Comment by simonw 9 hours ago
(This is a subtle nudge at anyone from OpenAI who reads this to make sure they get updated.)
OpenAI have a model called "chat-latest" - I wonder if that's running this new model yet: https://developers.openai.com/api/docs/models/chat-latest
It's described as "points to the latest Instant model currently used in ChatGPT" - so presumably that's "GPT-5.6 Instant" in the app.
Comment by firasd 8 hours ago
Comment by simonw 8 hours ago
https://gist.github.com/simonw/aae4febd3c6f7bc5b7811857edb3c... has screenshots that still show "Instant" as an option for ChatGPT Chat... but not for ChatGPT Work.
Comment by firasd 8 hours ago
It’s hard to understand this .. like sure we can select instant but is there an actual model called 5.6 instant? Like is it on LM Arena and OpenRouter or available via API etc
5.5 instant is definitely A Thing it’s even name checked in this OAI post
Comment by simonw 7 hours ago
Comment by hluska 3 hours ago
Comment by tosh 10 hours ago
luna is very good
Comment by timpera 9 hours ago
Comment by skybrian 9 hours ago
Comment by redox99 9 hours ago
Comment by stuartq 8 hours ago
Comment by Sammi 8 hours ago
Comment by LaurensBER 7 hours ago
Both models are cheap enough that I can run 4 sessions at the same time without running out of the 20 USD codex and 10 USD Opencode plan. I've burned through almost a billion tokens this week and I've done some pretty big refactors as well.
I have a Claude Max subscription but I've barely been using it because of the many issues they've had this week.
Comment by dannyw 3 hours ago
Pricing absolutely matters.
Comment by redox99 7 hours ago
Comment by causal 9 hours ago
Comment by msq22 9 hours ago
Comment by timpera 9 hours ago
Comment by kaszanka 3 hours ago
Hm, does this mean that 5.6 Pro in ChatGPT web is somehow different/not as good now? I found it really good for code review (upload your repo and patch and off it goes).
Comment by roytam87 2 hours ago
Comment by saithound 7 hours ago
The Aug 6 update has forced the entry box to auto-format Markdown in an attempt to imitate Claude. The implementation is buggy and even simple copy-and-paste has gone entirely haywire. They also forgot to leave a switch to turn the confounded autoformatting thing off.
Chat mode in general is currently crawling with more UX bugs than a porch screen in summer.
Comment by kgeist 7 hours ago
I wonder if they actually do it to optimize inference. I maintain a corporate AI server and one of the tricks to reduce the load was to modify the system prompt to be as terse as possible so the average response completes faster and requests queue up less often.
Comment by johnnyApplePRNG 3 hours ago
Yes, please give free users more access and leave paying codex users in the dust.
Wise plan, Sam.
Comment by sunaookami 8 hours ago
Comment by OsamaJaber 8 hours ago
Serving cheaply at that scale means routing, batching, and cache hits, not a better model :D
Comment by applfanboysbgon 7 hours ago
My mission is world conquest. I'm writing a comment on an HN thread.
No, those two clauses have no relation whatsoever. I just felt like saying the first sentence because it sounded cool.
Comment by Squarex 8 hours ago
Comment by HDBaseT 5 hours ago
Comment by ignoramous 10 hours ago
Every week, 1 billion people turn to ChatGPT for everything from quick questions and web searches to planning, research, advice, and complex decisions.
Guess, Google's AI Mode is chipping away at their consumers (I know I haven't used Chat in a long, long while for 'quick questions and web searches' after OpenAI did away with "think" which I always use). The money-minting office & coding market Anthropic has cornered is hyper-competitive at both the frontier & low-cost ends. OpenAI is reactive [0] and seems right up against it, despite the strength of its excellent models.[0] Won't put it past OpenAI (and/or Google) to open weight larger models!
Comment by skybrian 9 hours ago
Comment by drivebyhooting 9 hours ago
Comment by sk4rekr0w 2 hours ago
Comment by porridgeraisin 9 hours ago
Comment by iJohnDoe 55 minutes ago
I was not impressed with 5.6 and this hits exactly why.
Also, this instant, medium, and high slider situation we now have everywhere is batshit crazy. It’s a major step back in technology.
I’ll put money on the table there will be a surprise in revenue because users have no clue what to choose, so they are constantly choosing high because they don’t want to risk getting inaccurate answers. If this was intentional by the dark patterns department, then brilliant. However, I’m guessing Anthropic and OpenAI are struggling to know how to deploy their models.
Also, the models were already amazing. They need to slow down and do a model release once a year and only do extremely minor iterations instead. Some amazing things are accomplished, but how people actually want to utilize the models gets screwed up every time in the process.
Comment by simianwords 9 hours ago
Comment by redox99 9 hours ago
Comment by jauntywundrkind 10 hours ago
But seeing the graphic with the visual weather report: that makes me think that is not the goal at all. :)
Comment by laweijfmvo 9 hours ago
Even after identifying the 0% chance of rain, it still drags the conversation on and on and on
Comment by taikahessu 9 hours ago
Comment by sunaookami 8 hours ago
Comment by aniceperson 9 hours ago
Comment by porridgeraisin 9 hours ago