Gemini 3.8 Live and 3.8 Live Extended Thinking
Posted by leumon 11 hours ago
Comments
Comment by jeanbza 7 hours ago
I've been using Gemini to live chat in Afrikaans and do impromptu Afrikaans grammar lessons during my solo drives around town. It is phenomenal at speaking the language - like, it really shocks my family members when they hear it.
This is probably the most joy I get from any of my usages of LLMs/AIs. It's been really, really nice getting to speak my language regularly again. =)
So, I'm excited about this release and live chat getting better. I also hope the other frontier labs pick up niche languages like this as well so that I have more options.
Comment by LluisGerard 6 hours ago
I once asked it to summarize The Hobbit in Catalan to explain it to my daughter before sleep. I was expecting a lot of mistakes as I see regularly if I ask anything in my native language when using GPT or Claude, but it was surprisingly good. I was going just to kind of skim ahead and retell it my own way, but ended up almost saying it verbatim because it was good already.
She loves Zelda so I asked it to explain the story of Breath of The Wild keeping the original names, and to make it fun, etc.. I was surprised again. I did retell some bits in my own style and taste but it is very convincing.
I haven't tried Catalan on newer models like GTP-6 Astra or Fable tho. We have all these benchmarks based on software development, and AGI, etc.. but it would be cool to have some language benchmarks for different communities.
As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of these benchmarks were made in other languages, would the result be similar.
Comment by jaggederest 4 hours ago
I am not a Chinese speaker but my understanding is that all of the models have substantially different behavior in Chinese, to the degree that it's kind of like a second model. Would be interested in hearing more if anyone has direct experience.
Comment by mncharity 25 minutes ago
Relatedly, when chatting with frontier models about science education content design, I've found it very helpful to mix in Chinese education terms. In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop". And "estimation" (educational) in the US means one (dysfunctional:) thing, which similarly distracts. Perhaps if AIs become increasingly multilingual, but remain weak at deep conceptual reasoning, it may be fruitful to have multilingual thesauruses, to use language-associated cultural conceptual differences as a way to convey conceptual nuances with which LLMs otherwise struggle?
Comment by eru 16 minutes ago
For that particular attractor, even German or Spanish or so might help avoid it? Or perhaps even just using British English?
Comment by asdfman123 6 hours ago
Comment by abhgh 6 hours ago
Comment by asdfman123 6 hours ago
How did the early Roman Empire interact with Greek city states?
How are LLMs planning on learning new information on the fly without new context or retraining runs?
Why did mom leave?
You know, standard stuff.
Comment by throwup238 5 hours ago
Reinforcement learning by human feedback.
Comment by rr808 2 hours ago
Comment by tharkun__ 1 hour ago
It tells me there's a Costco 10km from me and I'm like, yeah let's go there, that's the one I meant! And then it sends me to one that's like 40 minutes from me all across town and I'm like WTF?!
Never mind asking Android Auto any question that's not driving related. It just doesn't understand me at all. Or even some driving related ones. "Find Alternative route". "Sorry, I can't help you with that". WTF?
Comment by bberrry 3 minutes ago
Comment by jonifico 3 hours ago
Comment by cm2012 5 hours ago
Comment by arnorhs 7 hours ago
Comment by trollbridge 2 hours ago
Comment by 3stacks 6 hours ago
As jy wil, ons kan saam praat op Discord :) maar my Afrikaans is sleg
Comment by ramijames 6 hours ago
Comment by written-beyond 5 hours ago
Comment by ramijames 5 hours ago
If you're really interested, shoot me an email.
Comment by SoftTalker 1 hour ago
Comment by geraldwhen 4 hours ago
Comment by basch 5 hours ago
It makes total sense it’s good at regurgitating the edge cases of language.
Comment by jimmySixDOF 6 hours ago
Comment by HDBaseT 6 hours ago
Comment by Havoc 9 hours ago
Copes well with thick accent, voices are pleasant and latency seems low.
Oh and I can actually use it on a workspace account - which for most of the recent releases was an account stuck in limbo. Not personal enough for personal offering, not enterprise enough for enterprise.
Well done G - will definitely be using this
Comment by Havoc 9 hours ago
And looks like one can trigger live mode via siri
Comment by rdtsc 10 hours ago
Comment by WarmWash 10 hours ago
I can't think of a better general purpose model than 3.8 flash right now. It also writes more naturally than the other big models too.
Comment by bahmboo 6 hours ago
Comment by rdtsc 8 hours ago
Comment by asdfman123 6 hours ago
Comment by rdtsc 2 hours ago
Comment by mchusma 9 hours ago
Comment by mapontosevenths 9 hours ago
However, if I want it to DO something then Gemini is in absolute last place. I don't trust it for anything more than renaming files that I don't care about very much or extracting data (though it's too expensive for data extraction at scale).
Comment by kridsdale1 6 hours ago
It’s for this exact kind of scenario where a random question pops in to my head.
Plus, it’s the most grounded by real live data of all the chatbots.
Silicon Valley people are majorly sleeping on Google Search AI Mode.
Comment by james2doyle 7 hours ago
Comment by alex1138 9 hours ago
People use text with LLMs but it's great to have a high fidelity "analyze this image"
Comment by plaidfuji 8 hours ago
I think they’ve made a shrewd move in focusing on search integration and everyday users (Gemini app) vs software power users. They have their corner and nobody is really competing with them, plus it feeds directly into their existing revenue stream.
Comment by Eridrus 4 hours ago
Maybe everybody is wrong about how valuable these companies will be, but atm it looks very dumb to not be competitive in coding.
Comment by dansquizsoft 50 minutes ago
Occam's Razor is overwhelmingly that they just don't have the organisational capability to capture this market. If they did then they absolutely would have.
Comment by Keyframe 7 hours ago
Google in a sense won (me over) like that. I also expected them to brute force their way into everything and dominate. This is how it played out though. Image generation is great as well, but ChatGPT one is more lenient on copyright and nannying - for example when my kid asks me to "take a photo of him and Sonic". Gemini cops out either because of the kid or Sonic, disappointing us both, but ChatGPT can be.. persuaded.
Comment by roncesvalles 36 minutes ago
Comment by smcleod 6 hours ago
Comment by yoz-y 6 hours ago
Comment by yieldcrv 5 hours ago
these are essentially monthly releases, the dated releases that Deepseek and Qwen do make more sense
Comment by cmrdporcupine 9 hours ago
Ask yourself, how does Google -- a company that famously does everything -- benefit from SWEs outside of Google having access to powerful coding models? They would just be competition.
Famously, Google just eventually discards almost all businesses that don't have the same fire hose of revenue that ads does. Selling coding plans isn't something they are going to want to do.
Google is clearly motivated to make better search and information finding tools and stuff that will ultimately drive users through their existing search/ads/youtube ecosystem. That's really why they're in Android, that's why they do Chrome. Everything else with them is a sideshow.
Google is also full of beancounters obsessed with data centre quota and resourcing. Even massively profitable ads projects have to justify and fight for it. (Source: used to work there).
I can't think of anything less resource & revenue sensible than providing outside parties access to your TPUs for the purpose of letting them write stuff which could just end up competing with you.
Yes, maybe as part of their cloud business, selling token access could be useful money. But I doubt they'd tune it for coding.
Comment by rdtsc 8 hours ago
Bragging rights to say they have a SOTA model. I guess that was more like Google of 10 years ago with moonshot projects. Nowadays, yeah, perhaps if it's not helping sell ads, it doesn't make sense.
Comment by geodel 5 hours ago
It would be utter stupidity if they keep competing with hyper agile frontier labs. They sensibly moving to big infra provider and that's good niche.
> .. perhaps if it's not helping sell ads,
Perhaps cloud business is also a thing which is making big revenue
Comment by mlmonkey 8 hours ago
Catch is, even Googlers internally do not have access to top-tier models. (or did not until recently, when apparently Claude was made accessible to the SWEs internally).
Comment by kridsdale1 6 hours ago
Funny enough, after trying it, I went back to G3.8
Comment by VirusNewbie 2 hours ago
Comment by brap 8 hours ago
Comment by _s_a_m_ 10 hours ago
Comment by nolok 10 hours ago
As opposed to what, them not having it and burning money that isn't their instead like openai and anthropic? At least Google is feeding itself instead of having to create a bubble to stay alive
Comment by verdverm 9 hours ago
Comment by runako 9 hours ago
Big rich companies take on debt for reasons that are sometimes inscrutable from the outside. Recently, they have been borrowing for ~5%, about a half point above what the US government gets for 10-year Treasuries.
Apple has been financing operations with debt for a number of years as part of a complex optimization plan.
No, Google is not broke.
Comment by verdverm 7 hours ago
Some expert wall street analysts discussing what they found and how they dissect things, have a healthy skepticism of Big Ai
Comment by msabalau 8 hours ago
If my quick search is correct, Google is sitting on a quarter trillion dollars in cash and marketable securities. They could keep doing the negative cash flow thing at this scale for another decade.
Comment by jasondigitized 5 hours ago
Comment by nolok 9 hours ago
Comment by jnwatson 9 hours ago
Comment by haberdasher 10 hours ago
Comment by chpatrick 9 hours ago
Comment by galkk 41 minutes ago
Comment by solenoid0937 21 minutes ago
Comment by Zsfe510asG 10 hours ago
Comment by phenomen 9 hours ago
Comment by thisgoodlife 7 hours ago
Comment by riddlemethat 7 hours ago
Comment by safog 7 hours ago
Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.
I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.
Comment by BenzeneDream 6 hours ago
Comment by qingcharles 49 minutes ago
Comment by vrosas 6 hours ago
Comment by taylorfinley 5 hours ago
Comment by xnx 4 hours ago
Comment by andai 4 hours ago
alias agy="agy --dangerously-skip-permissions"Comment by cute_boi 1 hour ago
Comment by VectorLock 3 hours ago
Comment by baq 7 hours ago
Comment by le-mark 6 hours ago
Comment by robotmay 5 hours ago
Comment by adventured 3 hours ago
Comment by staticman2 7 hours ago
Anecdotally I'd rate Gemini behind Claude and OpenAI models at fiction and I can't find any benchmarks showing Gemini is the clear winner at this task.
Comment by jakderrida 7 hours ago
I doubt they even intended it to be, but it seems like I kept going from resorting to 3.5-3.8 (over time) to realizing that Claude and GPT, while great at Python, will make rudimentary mistakes with R; even when they compose giant complicated R code.
Comment by sosrobahu 12 minutes ago
Comment by WarmWash 10 hours ago
I'm worried in their push to catch up on the SOTA front, it's going to lose that natural sounding touch it currently has.
Comment by alansaber 9 hours ago
Comment by lilbigdoot 8 hours ago
Comment by porridgeraisin 8 hours ago
Comment by amelius 6 hours ago
My ChatGPT env only says "low", "medium", "high".
Is this a "pro" thing? I have totally no idea what I'm talking to, so actually I'm thinking of stopping my plan. Gemini and Claude are much more clear about it.
Anyway, I like the speed at which Gemini responds so indeed for simple things it is preferable.
Comment by comex 6 hours ago
I’ve been using Work for all my queries, since it seems to just be the same interface as Chat but with more features. I don’t understand why they’re two separate things.
Comment by willy_k 6 hours ago
Comment by lxgr 8 hours ago
Comment by pkulak 8 hours ago
Comment by kylecazar 7 hours ago
Their AI leadership team has taken some hits recently too, in the form of departures. I believe when they get their bearings they will be competitive again. 3.8 Flash has been a great model for me.
Comment by spwa4 7 hours ago
Comment by drivebyhooting 9 hours ago
Meanwhile Claude and Astra like to couch all their agreements with caveats and provisos.
Comment by re5i5tor 8 hours ago
Comment by indymike 6 hours ago
Comment by mapontosevenths 9 hours ago
Sometimes that's what being smart sounds like.
Comment by Someone1234 8 hours ago
Someone confident but incorrect, can often sound more convincing than someone with actual expertise. The expert must add caveats/hedge, because those are the facts on the ground, whereas the person reciting google can be entirely confident.
Of course the people judging aren't experts, so they side with confidence and simplicity. Heck, just writing shorter replies on Reddit is rewarded. Nobody reads the articles, let alone a paragraph-long reply.
That all being said though, there are limits. Sometimes LLMs on high-thinking go off on full tangents based on little, and don't have the self-awareness to bring it back.
Comment by JW_00000 7 hours ago
Comment by Someone1234 4 hours ago
So I won't be addressing this, for those reasons and others.
Comment by nostrebored 7 hours ago
Comment by mapontosevenths 6 hours ago
To an expert communicating with a layperson is a form of compression. You must turn some very complex idea into one that you suppose the other person can grasp given their limited frame of reference. It's always lossy, and you have to guess how much you can remove without sounding patronizing or being inaccurate. It's tough, and the more you know the tougher it gets.
Ever done that "explain what happens when I visit Google in my web browser" interview question?
A sales guy will answer in a sentence. An engineer might be able to talk about it for several days and still not be sure they didn't miss anything important. That much knowledge can actually be detrimental to communication.
Comment by drivebyhooting 9 hours ago
Comment by copperx 9 hours ago
Comment by jakderrida 7 hours ago
When you're a ChatGPT Projects or Claude Projects user, those caveats and provisos are your worst enemy because they'll change caveats into hard rules (either for the session or committed to memories) and you end up in absolute hell having to make it investigate to figure out why it can no longer produce anything but read-only pre-check code that never actually does anything but keeps performing stupid safety checks.
Comment by avereveard 6 hours ago
Comment by ankurshv 2 hours ago
Comment by porridgeraisin 8 hours ago
Comment by chicagobuss 6 hours ago
Comment by martythemaniak 9 hours ago
Comment by NBJack 8 hours ago
Comment by avereveard 6 hours ago
Comment by robotmay 5 hours ago
Comment by hiimkeks 8 hours ago
> This mirrors how Apple has always segmented Pro vs. non-Pro iPhones: base models got LTPS panels while Pro models got LTPO, and only with the mainline iPhone 17/17 Plus did that gap close the standard versions previously lacked the smoother 120Hz ProMotion technology and the always-on display feature, unlike the Pro models — the 17e is the one model line still using the older, cheaper panel.
(emphasis mine)
I mean, I can guess what it is trying to say, but who RL'd this nonsense?
Comment by malshe 8 hours ago
Comment by asdfman123 6 hours ago
It reminds me of a pedantic grad student.
Comment by YoumuChan 1 hour ago
Comment by schainks 9 hours ago
Comment by BeetleB 8 hours ago
Only worked in a 1:1 in a quiet place. Still, can't complain for free.
Comment by bahmboo 6 hours ago
Comment by BeetleB 6 hours ago
Comment by bahmboo 5 hours ago
Comment by tonyhart7 8 hours ago
well then its not model problem
Comment by BeetleB 6 hours ago
Comment by nutjob2 7 hours ago
Comment by burky 2 hours ago
Comment by bilsbie 8 hours ago
Comment by spider-mario 7 hours ago
Comment by chpatrick 9 hours ago
Comment by ghoshbishakh 9 hours ago
Comment by gunalx 8 hours ago
Comment by throwaw12 9 hours ago
Lately it became load-bearingly-reality-difficult to not only read, but to comprehend the Claude output
Comment by cromka 9 hours ago
Comment by Yizahi 6 hours ago
Comment by flyinglizard 9 hours ago
On my TODO is try and run all of the analysis pipeline in dense "machine speak" to save on tokens and just let Gemini sort it out at the end.
Comment by ympb121 9 hours ago
Comment by iamjackg 9 hours ago
Comment by saurik 9 hours ago
Comment by saurik 7 hours ago
And like, it does this despite it speaking in extremely dense math, which both makes it sound correct and requires a lot more effort to prove when it is wrong... yet, it isn't actually correct more often, and so that time sink just isn't worth the benefit. I then think many people--including people who can speak math (as can I)--just stop bothering to check everything, as if you come across a human who speaks like this it probably does correlate with slow and careful thought that helps prevent errors.
Instead, Claude has the mistake rate of a somewhat accelerated beginner impossibly combined with the language of an expert professor; and we as humans just aren't good at that combination: it becomes very dangerous and makes it take longer to spot its egregious mistakes and trained-in biases. If you have to use Claude, I thereby claim you really need to have a team of not-Claudes to help insulate you from this, and Gemini (while being a bit senile) is a lot more collaborative and approaches problems in ways that makes it harder to get tricked.
(To translate this into more of an engineering analogy: Claude always feels to me like the engineer who put more effort into learning how to program in functional languages than into how to actually develop working code, and then confidently presents you answers in Haskell or Lisp that never quite work. To find their errors is then very costly. In contrast, Gemini feels more like a Java or Go developer who knows they are a cog... that's helpful! <- Which maybe just goes to show that AI has finally turned me into a manager, omg.)
Comment by le-mark 5 hours ago
Comment by dlss 9 hours ago
Comment by paradox460 7 hours ago
I've set my documentation sub agent to Gemini and my code agent to Luna
Comment by deviation 10 hours ago
Comment by wongarsu 8 hours ago
Comment by time0ut 6 hours ago
Comment by yipinwong 5 hours ago
I am not so advanced enough as a human being.
Comment by monroewalker 10 hours ago
Comment by incompressible 2 hours ago
Comment by laweijfmvo 9 hours ago
Comment by stranded22 8 hours ago
Comment by TomGarden 8 hours ago
Comment by doodlesdev 10 hours ago
Excited to try this out! Shame on Google for not releasing Gemini 3.8 for Google AI Plus users yet, though.
Comment by ilaksh 9 hours ago
Comment by SyneRyder 7 hours ago
Comment by coda_ 2 hours ago
I'm not very happy with live conversations with Claude, so this seems like it might be a good option.
Comment by sahaskatta 10 hours ago
Comment by phenomen 9 hours ago
Comment by sahaskatta 9 hours ago
Comment by xd1936 9 hours ago
Comment by cnobody 10 hours ago
Comment by giancarlostoro 8 hours ago
Comment by verdverm 1 hour ago
Comment by johnsmith1840 7 hours ago
Is this a pure TPU infra? Really high performance solid intelligence.
Comment by rrr_oh_man 1 hour ago
Comment by laichzeit0 1 hour ago
Comment by samuelknight 9 hours ago
Comment by TomGarden 8 hours ago
Comment by bengkoang 3 hours ago
Comment by svcrunch 5 hours ago
Comment by qudat 7 hours ago
Comment by chrystianpl 8 hours ago
Comment by jeffbee 8 hours ago
Comment by smithcoin 9 hours ago
Comment by tonyhart7 8 hours ago
Comment by farnulfo 49 minutes ago
Comment by system2 1 hour ago
Comment by ghoshbishakh 9 hours ago
Comment by qlte 8 hours ago
There's one or two I find more understated but I would love a 2026 SOTA V2V model that speaks clearly but without the artificial personality layered on.
Human interaction/theory of mind relies so much on non-verbal clues for interpreting emotion/intent and so for me having those neurons firing constantly while talking to an LLM just for an emotional no-op is exhausting to put up with for more than a couple minutes.
There's one male voice that would make me assume someone was sarcastically mocking me if I was talking to an actual person because it's just so over the top.
Comment by jiggawatts 7 hours ago
I don’t want to be aroused by my turn-by-turn street directions, thanks.
Comment by mvdtnz 9 hours ago
Comment by nharada 9 hours ago
Comment by parasti 50 minutes ago
Comment by lostmsu 9 hours ago
Let me tell you unlike every other mentioned model Gemini 3.8 Flash trial had to be reverted the same day. Instead of simply delegating tasks it would invent additional requirements and implementation details it knew nothing about and no amount of convincing not to do it would work. That's the first time a model failed on me so spectacularly despite having practically same Artificial Analysis Intelligence Index as another model that just worked (and higher than working DS Flash).
The reason I think it is relevant is: Live is likely even stupider model in every way possible (except hearing better than separate STT). So beware using it for agentic scenarios.
Comment by blovescoffee 10 hours ago
Comment by tantalor 9 hours ago
Comment by toddmorey 7 hours ago
Comment by DonsDiscountGas 4 hours ago
Comment by attels33 10 hours ago
Comment by water-drummer 1 hour ago
Comment by verdverm 9 hours ago
Comment by SomeonesAccount 10 hours ago
Comment by bilarikan 9 hours ago
Comment by glimshe 10 hours ago
However, the "Extended Thinking" should be renamed to "Slightly Extended Thinking". Considering that it's the maximum thinking option for Gemini Flash in the chat UI, it doesn't actually think a whole lot, leading to an uncomfortably high number of incorrect/poor replies.
Comment by addandsubtract 2 hours ago
Comment by tiahura 10 hours ago
All audio generated by our AI products is watermarked with SynthID. This imperceptible watermark is woven directly into the audio output, ensuring AI-generated content remains detectable to help prevent misinformation. For details on our approach to safety and responsibility, review the model card.
Comment by hajile 9 hours ago
Comment by scottchiefbaker 5 hours ago
Comment by 740273730191 9 hours ago
Comment by SSLy 6 hours ago
Comment by verdverm 9 hours ago
Comment by varispeed 10 hours ago
Comment by andrewinardeer 7 hours ago
Looking forward to where this can go.
Comment by bronlund 10 hours ago
As PrimeTime said; these are the guys that invented the 'T' in 'GPT', that deployed their first TPU in 2015, that is using billions on AI - and they are beaten by 300 people startup named Moonshot AI even. People are going to write books about this complete fumble.
Comment by arw0n 9 hours ago
And Gemini is kinda good enough at everything. Never the top, but it is decent at every task, and it is much faster than Kimi K3 and significantly cheaper. Kimi is very focussed on coding, Gemini isn't.
More importantly, it natively understands text, audio and video. If/when we are able to make the jump to robotics, this becomes essential. As you say, Google has a lot of deep background and deep pockets, they are able to make more of a long play. No idea if it will pay off, but it is way to early in the game to count them out.
Comment by bronlund 9 hours ago
Comment by verdverm 1 hour ago
Comment by password54321 9 hours ago
My advice is to listen less to brainrot 'influencers' that optimise for engagement through sensationalism.
Comment by bronlund 9 hours ago
They have "unlimited" resources and has researched AI since the very beginning - PageRank is a form of AI even. And still, Gemini is behind Claude, GPT, Grok, Muse, GLM, Kimi and is maybe on par with DeepSeek?
As I said, it is embarrassing.
Comment by password54321 9 hours ago
Comment by Forgeties79 4 hours ago
Comment by no_carrier 2 hours ago
Comment by Forgeties79 9 hours ago
No one is behind grok. It literally has "be funny and irreverent when appropriate" (whatever the hell "when appropriate" means for them) baked into the system prompt. To me, that is all you need to know about how useful it is.
No serious people use it and the numbers bear it out tbh. It has the smallest market share of the "big companies" for a reason - and it's by a very, very large margin (~2.5% last I checked).
Comment by mattlondon 9 hours ago
And they're making money doing it.
Perhaps they don't have the best coding model right now (although 3.8 flash is arguably SOTA at some benchmarks), but is that the be-all and end-all of AI? Only coding matters?
Comment by verdverm 1 hour ago
Just because it shows up does not mean it is being used. I only click the feedback button to tell them how much I dislike their Ai summaries. They have removed that feedback button this week. Can't take the heat I suppose.
I've also never know anyone who finds them trustworthy nor heard someone say anything besides how they also dislike them
Comment by c0nducktr 42 minutes ago
"It's AI but anyway", seems to be a way to use it, but not promote that your using it?
There's a significant amount of people in my friend group who don't like "AI", but everyone seems to use what google's doing, even if they add the "the AI said this" disclaimer to it.
Even the people who don't like 'AI', will still reference the LLM output, which seems like they're just slow to accept it, I guess.
Comment by thereitgoes456 9 hours ago
Comment by dbbk 9 hours ago
Comment by owebmaster 1 hour ago
Comment by verdverm 1 hour ago
Comment by brazukadev 47 minutes ago
Comment by polski-g 2 hours ago
Comment by verdverm 58 minutes ago
Comment by mchusma 9 hours ago
Comment by Dardalus 6 hours ago
Comment by dbbk 9 hours ago
Comment by dude250711 9 hours ago
Comment by fileeditview 10 hours ago
Comment by tonfa 9 hours ago
Given the very high margins on inference, once volume is large enough the other can also start printing enough money.