Ask HN: What default model do you use and why?
Posted by stikit 3 days ago
I use claude for most of what I do, and Fable is largely overkill for me and frequently burns through my Max plan's session credits in minutes (!) when just doing an initial mobile app planning with 4 agents. After I waited out the timeout period 6 hours later, and I picked up again, the cache had timed out so it burned through 2% of the session in less than a minute. Opus 4.8 is now my goto and I will be avoiding 5 until I see a reason to switch back. 4.8 is 'good enough' for what I need and has been a great value. It mostly gets things right. Most of what I do is web and mobile , largely cloud backend.
Comments
Comment by o_m 3 days ago
I also don't want to use the Claude Code and Codex agent harnesses. The good thing with Codex subscription is that it can be used in other harnesses, unlike Claude. As far as I know, only Anthropic has this restriction.
Comment by pid0x17 5 hours ago
It generates pages for what could be explained in 1-2 paragraphs.
I use to love using Gemini CLI to have Gemini explain a module for me before I dive into the code and point out specific points of interest, but I can't do that with Claude Code without feeling overwhelmed and starting to heavily procrastinate
Comment by r_lee 3 days ago
It's really strange because when Opus 5 released, there were some that pointed this out, but a bunch simply said it was the best and as good as Fable etc etc.
but for me, it caused me to get very demotivated and avoid interacting with the model, at least when using Claude Code.
Comment by Inozem 3 days ago
Now I’ve gone back to Codex because I simply find the ChatGPT models much more comfortable to work with. I spend less time fighting the model, correcting its direction, or re-explaining what I meant.
Comment by wenc 3 days ago
I've moved to Codex 5.6-Sol. Much saner English, much better at execution, and gets stuff done in a matter-of-factly kind of way (Claude Code is a mess these days -- it gets things wrong and goes around in circles).
But I'm harness agnostic and am not locked in. I just keep my issues in Kata Tracker (https://www.katatracker.com/) and switch harnesses/model when I need to.
Being loyal to a particular model/harness seem unwise to me.
Comment by amelius 3 days ago
I suspect this degradation is happening because the AI labs are using the LLM's output to feedback into the input, to create a thinking loop, and they're optimizing that.
Comment by montroser 3 days ago
For what it's worth, here's a take on its speed vs cost vs intelligence: https://artificialanalysis.ai/models/deepseek-v4-1-flash
I can go all day and night with this thing with multiple sessions going, and I spend like $2 per day retail. With opencode-go, that fits within the $10/mo subscription, so that's what it ends up costing in real life.
Comment by the__alchemist 3 days ago
Comment by montroser 3 days ago
In cost-per-task, it's a 32x multiple Fable vs DS41. Either it really does cost them more electricity/gpu, or else they are succeeding convincing the masses that Anthropic's models are necessary for everyday tasks, when this is truly just not the case. Either way seems a tragedy.
Especially regrettable from my perspective, is the idea that to do computer programming in these modern times, you have to fork over an amount comparable to many people's monthly rent. Even if your employer is happy to direct their funds that way, we lose quite something when recreational coding comes with a substantial personal monetary cost.
Comment by cassianoleal 3 days ago
Comment by prettyblocks 3 days ago
Comment by cassianoleal 3 days ago
Comment by Obscurity4340 2 days ago
Comment by the__alchemist 3 days ago
Comment by AaronAPU 3 days ago
Comment by montroser 3 days ago
I'm glad that you have the luxury burn $200/mo. But that aside, it's not obvious to me that handing your pennies to an American company that is in cahoots with the American military, willing to self-censor at the direction of the American government, investing in large-scale raping of the planet to build new power plants data centers -- is better than handing your (20x fewer) pennies to a Chinese company who releases its models in the open, and is investing in efficiency so that we can run the same inference for a fraction of the resources.
Comment by Fizz43 2 days ago
Comment by montroser 2 days ago
Comment by AaronAPU 3 days ago
Comment by ricardobeat 2 days ago
Comment by hgoel 3 days ago
Used to pay for a Claude 20x plan and did everything in Opus, but I hate how it talks now and recent events (OAI scooping, Anthropic's spying, third party Chinese model hosts stealing and selling credentials) have really pushed me towards local AI for personal needs. Am not allowed to use Chinese models for work even if self-hosted so not much choice there.
Comment by oidar 3 days ago
I'd love to hear more about your setup. I have a single GB10 and am thinking about adding an additional one.
Comment by hgoel 3 days ago
Initial setup was a tad annoying because I had to update their firmwares and then power cycle them to get the 200GbE link working at full speed. After setting that up, it has been pretty smooth. I don't directly deal with the cluster, usually I just have the LLM itself handle updates/stopping to load different models.
Generation speed and TTFT is decent with Qwen3.8-flash, and it does a good job for my fiddling around with enough concurrency for multiple sessions/subagents. GLM 5.3-flash was also nice, but not too much better for how much slower it is.
I should also add that I already maintain a homelab with a couple of computers, VMs etc, so I am probably somewhat more tolerant of the occasional issue and fine with manually managing stuff over SSH. I think this is just a tradeoff of self-hosting relatively recent tech though.
I have a triple 3090 rig, but it mostly stays powered off because of the massive power draw and cooling requirements. The Sparks are slower but at peak they consume as much power as my 3090 machine at idle.
The recent talk of regulation has me wanting to pick up 2 more Sparks, but that's mostly to have the capacity to play with multiple models, local model tuning and to be ahead in case they force some limits/registration requirements for buying new hardware (kind of like the attempts to regulate 3d printers).
Comment by tmikaeld 2 days ago
Comment by hgoel 2 days ago
What I get out of it is the ability to hand login credentials to my other computers to manage their updates, bug fixes etc. Eg. After updating my proxmox server, the nvme drive kept dying. Was able to let my local AI in to figure out and fix what was wrong (known issue). A cloud-based AI could've done it too, but I don't want to be sending internal passwords out of my network like that.
Plus, the ability to freely delegate tasks or exploration of things cloud models generally avoid. For example, I draw as a hobby, and when I'm struggling with a pose but can't quite figure out what I'm missing, I pass it into a VLM for advice, but Claude etc get unnecessarily cautious because they interpret an anatomical sketch as a naked person.
Comment by vallerie 3 days ago
- like a fancy auto complete (here are some stub methods, they should do X, fill them in)
- using fairly detailed plans and test harnesses, so blowing up the world is hard
The 3.X Flash family have been fairly capable models, and the selling point for me is just raw speed. Gemini is noticeably faster than the competition, about 3-4x, and I just get work done faster with it.
That said I'm keeping an eye on Open Weights. DS4 Flash was good until price hikes, and finding a provider that serves at high speed and without quantisation at the prior price is tricky.
Comment by martythemaniak 3 days ago
There's no Gemini pro model currently, so you gotta pair that with a 20 openai plan for access to more advanced stuff if you need it.
Comment by vallerie 3 days ago
A few months ago that would've been bigger model stuff, but now you can do that in a couple hours with Flash and direction.
Comment by asd88 3 days ago
Astra’s outputs are much more concise while being similarly accurate, they seem more information dense. It’s also extremely fast and doesn’t have to “go check instead of answering from memory” if the answer is already in context, or dig super deep into a project before answering.
It’s honestly refreshing to work with Astra after being basically burned out from reading Claude’s responses.
It’s not perfect. It’s just as “mid” for professional SWE work as Opus/Fable, regularly making incorrect assumptions/generalizations, requiring steering in large codebases, and being incapable of making reasonable long-term software design decisions on its own in complex projects.
Comment by ulrikrasmussen 3 days ago
Comment by beej71 3 days ago
Comment by hmokiguess 3 days ago
Comment by seanmcdirmid 3 days ago
Comment by time0ut 3 days ago
I prefer Cursor at this point just because of Composer. Claude Code is passable but the lack of a good, fast, cheap workhorse sucks. Sonnet and Haiku aren’t it.
I also do like Codex and Sol, Terra, and Luna. They are decent but I don’t find they stand out enough to use over the others.
Additionally, I have not tried Astra and found Fable to really not worth the cost for the tasks I do.
Finally, the latest Grok is actually a beast of a model, but expensive enough to not be a stand out.
Comment by m0rde 3 days ago
Defaulting to Sol Medium/Light for most planning and implementation I think is tricky or want more care in.
Luna Extra High for everything else (implementation, tedious take over my browser and do stuff).
I read all of its output tokens and lots of thinking tokens to understand the general flow of things, but only minimally look at code these days. I can't grok what Claude models speak and it's gotten worse. OAI models speak my kind of tech language I guess.
Light human review, some automated review.
Most work is for internal use.
Comment by codazoda 3 days ago
For my personal stuff, I'm on a small $20 plan, so I need to use tokens conservatively. I was very rarely exceeding limits until I built a Dark Software Factory. It's not as efficient at token use. So, I use Sonnit over Opus here.
At work I have a $100 plan that I rarely exceed so I use Opus. I have access to Fable too, and I did use it a lot while it was new, but I don't find it improves most of my work by too much. I do mostly bug fixing across several hundred repositories with hundreds of thousands of lines of code, mostly written by humans over the past 20-years. These projects interact with each other so I run claude from the root of my sandbox (I was nervous to try this but I'm not looking back now).
I also use Sol as a secondary for my personal work. I pay for it because I like to talk to ChatGPT on the web. Since I already have the subscription, I let Sol write plans for me. It does a better job at certain tasks and it saves me some Claude tokens. Maybe I should consider Terra for the task, but I don't run up against my usage limits for the little bit I use it.
I'm trying to use Gemma 4 12b for some workloads but I haven't mastered the model yet. It's still very experimental for me. I can get work from it but it takes a lot of hand-holding. For local models, however, it's all I have the RAM for.
Comment by pighive 3 days ago
Comment by codazoda 2 days ago
https://joeldare.com/creating-a-minimal-dark-factory
So far I’ve created a couple simple web based games and a BASIC32 interpreter that boots on an ESP32 and turns it into a 1980’s style computer.
Comment by joelboersma 3 days ago
Outside of work, I'm using Codex's free tier with 5.6 Luna-High. My personal/volunteer projects are simple enough that I don't need much more.
Comment by Jordan-117 2 days ago
I try the ChatGPT free tier as a backup second opinion sometimes, but find their newest model to be painfully rambling.
I want to like Claude and might consider subscribing to it instead, but am put off by reports of its tight usage limits. With Gemini, though, I never hit a limit unless I conduct a few Deep Research queries at the highest level.
Comment by jakevoytko 3 days ago
As a side project, I'm doing an experimental task now (having Astra implement my own personal Google Docs clone, OT and everything, for my blog), and it feels more capable than Fable but needs to be watched closer than Fable. It is obsessed with verification and evidence to the point that I've needed to make some absolute rules to stop it from repeatedly e.g. making unit tests with 10,000 test cases with 5+ hours of runtime
For well-defined coding tasks, I'd default to either Luna or Sonnet, or hell, just writing it by hand.
Comment by george_stephen 3 days ago
Grok is keeping up and really good at explicit stuff, i think that's a drawdown for me. I think most nsfw was made with grok
As for meta ai, i feek it's a total joke, i don't have access to muse cause meta has decided not to support older OS versions (I'm on Android).
This comment box makes me feel almost as if I'm coding, because of the font. Pretty noice
Comment by gtadesktop1 3 days ago
Comment by nvme0n1p1 3 days ago
Comment by gtadesktop1 3 days ago
Comment by alstonite 3 days ago
Comment by ewindisch 3 days ago
Astra low on the $200/mo "20x Pro" plan gets me through a single day.
Comment by jbonatakis 3 days ago
Comment by fibonacci112358 3 days ago
Comment by ewindisch 2 days ago
>300 merged PRs/week
Comment by expedited123 3 days ago
Comment by wdm0006 2 days ago
Comment by conradludgate 3 days ago
For work, given we pay API pricing anyway, I've been happy experimenting with K3-medium as my default, and now I'm planning on trying GLM 5.3 as well. Sol-medium is my fallback for my work purposes but I occasionally use Opus 4.8 as an additional reviewer.
I like the open weight models because much of my time at work is spent on security hardening (specifically, hardening my own service), and Sol/Opus keep snitching on me and blocking my prompts.
Comment by gpugreg 3 days ago
GLM-5.3, GLM-5.3-Flash and Kimi K3 are also fine, but slower, more expensive, and less good for what I use them for, which is mostly Python and CUDA programming with some JS and HTML inbetween.
But almost all of my tasks are verifiable tasks, which means that the LLM can check whether it is done or whether it needs to keep trying. If you are mostly working on problems where the quality metric is based on vibes, YMMV.
I haven't tried Astra or Fable yet, because I am not made of money and am happy with my current setup. Also, the Opus models' writing is absolutely insufferable. My blood pressure rises every time I see a Claude-generated slop README.
Comment by mariocesar 3 days ago
For more long work, I now use Fable to create a PLAN.md. I tell it to make a plan that will be executed by other models, and most of the time it ends up choosing Opus or Sonnet.
I didn't start doing this recently. Before that, I would just use the top model for everything. Splitting the work across different models depending on the task has helped a lot. They run faster, and I usually get much better results
Comment by sourcecodeplz 3 days ago
Comment by mariocesar 3 days ago
Here is the script https://github.com/mariocesar/dotfiles/blob/main/common/.loc...
Comment by notnmeyer 3 days ago
Comment by oduis 3 days ago
Comment by the__alchemist 3 days ago
Granted, yesterday I threw a few tasks to Astra which the former 2 botches; it produced clean, correct solutions quickly, so pending further eval, this may take over.
IMO unless it's a mechanical tasks, it's worth it to use carefully -crafted queries on the more expensive models, than iterate through messier solutions on the cheaper ones.
edit: maybe not. Astra has imitated Fable, and appears unusable for biology due to safeguards.
Comment by verisimilidude 3 days ago
For high-level strategic conversations, I use Fable.
For planning, I use Opus or Sol. Sol is generally preferred; it’s faster, cheaper, and less verbose. But I still find Opus more capable on the most nuanced or complex tasks.
For planned implementation, I use Sonnet.
For one-shot unplanned implementation, I’ll use whatever model seems best for the task. FWIW, I’m increasingly turning to Grok here.
I use Luna all over my workflow for reporting.
Comment by traverseda 3 days ago
Comment by karmakaze 3 days ago
At work mostly Opus 4.8 (sometimes a GPT or Gemini 3.1 Pro). I find Opus 5 chatty/slower and Fable can venture into over-engineering itself into unnecessary complications.
Comment by ricardobeat 2 days ago
My current go-tos are DS Flash, Minimax M3 for subagents. Adding Muse Spark 1.3 to that mix but not convinced yet.
Comment by sourcecodeplz 3 days ago
unbeatable price/intel ratio per M tokens:
$0.10 (input)
$0.20 (output)
$0.002 (cached-input)
Comment by shelled 3 days ago
A day ago I activated Google AI Pro free via Google's tie-up with a local company (I do pay for this company's product though and it's anything but costly). Now I will use this too.
No other reasons to pick these, or not picking anything else.
Comment by seabrookmx 3 days ago
I'm not as up to date on the other vendors' models, but when I last used Gemini my feelings were similar between Pro and Flash.
Comment by jinnko 3 days ago
Comment by prism56 3 days ago
V4 flash latest has been great for me. It's so cheap it's almost free.
Comment by rsyring 3 days ago
- LLM expense budget
- What type of dev: work, personal, real time spaceship thrust vectoring, html contact forms for family, etc.
- human in the loop with short as possible turns, software factories that can run for days, or something in the middle
Just off the top of my head. I'm sure there are others I'm missing.
Comment by amelius 3 days ago
Comment by tiluha 3 days ago
Privacy policy is not great with deepseek api, but you are always just trusting their word with any hosted llm and in theory i could at least self host the models i use from deepseek.
Comment by donatj 3 days ago
Comment by herpdyderp 3 days ago
- Preferred: Claude Code with Opus 5 Medium
Comment by rpmisms 3 days ago
Comment by bellowsgulch 3 days ago
most engineering tasks don’t require frontier llms
when they get stuck, then i consider moving up to more capable models
purchasing a claude plan seems widely unnecessary to me
the tasks they do better than the average engineer cut both ways: unless you have an existing portfolio of well written and designed work done pre-llms, it looks like you’re producing slop that pretends to be well designed
poor typography choices despite using the mode,
poor layout choices despite using popular CSS frameworks
etc
bad engineers will always be bad engineers
tools don’t make up for it
edit: a follow up to this— everyone is using eyebrows in their layouts and have no fucking clue why it was done to begin with
everyone has a status pill floating above their front page hero display text and its not fucking status related
so gross
Comment by admiralrohan 3 days ago
Comment by bellowsgulch 3 days ago
Comment by NishanStepak 2 days ago
Comment by philbo 3 days ago
Comment by securekomodo 1 day ago
Comment by ewindisch 3 days ago
Comment by sourcecodeplz 3 days ago
Comment by adar2378 3 days ago
Comment by ghosty141 3 days ago
Comment by shamsalom94 2 days ago
Comment by guilhas 2 days ago
Comment by behole 3 days ago
Comment by nickthegreek 3 days ago
i find that luna can work through a 5hr quota window on a /goal without hitting 0%.
Comment by sidibe 3 days ago
Comment by jamesponddotco 3 days ago
If I don’t care about the code, i.e. I’m writing something quick just to test something, I’ll use GPT Astra to save my Fable tokens. For conversations and research I usually go with GPT Astra Pro.
For Home Assistant I use a combination of Grok 4.20 and DeepSeek Flash 4.1. Grok as a voice assistant, because it’s the only model with good latency here in Brazil (it’s nearly instant), and DeepSeek for everything else.
I don’t really use LLMs for writing, but I’m writing a fiction book on the side, and when I tried to use one for ideas, Gemini Flash 3.8 gave me the closest to good writing out of the ones I tested. Not enough for me to use it for the task, though.
Comment by Apreche 3 days ago
Comment by relug 3 days ago
Comment by JodieBenitez 3 days ago
Comment by w22oop 3 days ago
Comment by rajay99 1 day ago
Comment by Mohamed_Amineio 3 days ago
Comment by VariousPrograms 3 days ago
Comment by Cakez0r 3 days ago
Comment by jraedisch 3 days ago
Comment by cyanydeez 3 days ago
Comment by Havoc 3 days ago
...and then sprinkle in some other models when i think a second opinion will help
Comment by slipwalker 1 day ago
Comment by freeourdays 1 day ago
Comment by heypicoai 1 day ago
Comment by crusbro 2 days ago
Comment by rajeasy 2 days ago
Comment by mytime2 3 days ago
Comment by sebastienburel 2 days ago
Comment by maorbril 3 days ago
Comment by Inozem 3 days ago
Comment by irusik 2 days ago
Comment by fermataflow 2 days ago