DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
Posted by theanonymousone 3 days ago
Comments
Comment by pmxi 3 days ago
https://files.parasmittal.com/openai_aa_luna_dsflash.svg
1: https://openai.com/index/advancing-the-price-performance-fro...
Comment by fcanesin 2 days ago
High hopes for V4 Pro
Comment by stavros 1 day ago
Comment by k__ 2 days ago
Comment by bwfan123 3 days ago
So, are they planning to announce an optimized coding agent harness as well ? DSv4 flash is a fantastic model, and my daily driver. With reasonix or pi, I can code all day long and pay a few pennies for it. No token anxiety. Whereas the same model with fireworks/openrouter, with zdr thrown in, token costs ratchet up with no explanation. Likely that the model is subsidized for gathering usage data. I am waiting for the day I can run this locally.
Comment by segmondy 3 days ago
"For the Code Agent tasks among the public benchmarks above, DeepSeek-V4-Flash-0731 is evaluated with the minimal mode of DeepSeek Harness (to be released) as the agent framework, using the max reasoning effort level with temperature = 1.0, top_p = 0.95."
Comment by bwfan123 3 days ago
I meant something I could download and run.
Comment by ignoramous 3 days ago
Comment by jimmydoe 3 days ago
Comment by L8D 3 days ago
Comment by bwfan123 3 days ago
I am using openrouter with zdr guardrail which routes to any provider that supposedly doesnt train on user data. I also use fireworks (directly, not via openrouter) which is a provider promising zdr and has a bunch of open weights models. My issue is that these zdr providers dont transparently disclose caching/tokens etc and so they end up being far more expensive than directly using DS.
Comment by x312 3 days ago
Comment by bwfan123 2 days ago
Comment by smartbit 2 days ago
OpenRouter doesn’t list DeepSeek as ZDR combined with Caching. Thanks for the tip of Fireworks as Fireworks has Zero Data Retention by default. [1]
[0] https://openrouter.ai/docs/guides/features/zdr
[1] https://docs.fireworks.ai/guides/security_compliance/data_ha...
Comment by 0cf8612b2e1e 3 days ago
Does the file hosting actually cost peanuts when you do it yourself and the cloud has shattered my understanding of what it actually costs to deliver so much data?
Comment by wongarsu 3 days ago
At the scale of Huggingface, that still amounts to a lot load of money. Significantly less than if you did the same in AWS, but still a lot
That said, they do have a deal with AWS to make the data available in AWS ip space. Maybe they got some cheap hosting out of that too
Comment by oceanplexian 3 days ago
Comment by flyingpenguin 2 days ago
Comment by Bayart 2 days ago
Comment by gandreani 3 days ago
Comment by Bayart 2 days ago
As long as you're smart on caching on the edges and deduplication/overlaying with your content (which structured data types like models, containers and repos typically are fit for) you can get remarkably far for less than you'd think.
The economics of serving models seem to be way more dodgy.
Comment by anon373839 2 days ago
Comment by lima 3 days ago
Comment by himata4113 3 days ago
Comment by Scoundreller 2 days ago
Comment by miyuru 3 days ago
Comment by jvuygbbkuurx 3 days ago
Comment by rupx 3 days ago
CDN costs are pennies compared to inference and training though, HuggingFace will just get another 100 million and be set.
Comment by throwaw12 3 days ago
Comment by jmathai 3 days ago
It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it.
Comment by weiliddat 3 days ago
Comment by jmathai 3 days ago
One feature of the app is that all scripture is verified and what’s show to the user doesn’t come from the LLM at all and instead a trusted source.
I think exploring scripture this way does not alleviate you from struggling to learn and apply it. It hasn’t for me.
Comment by weiliddat 3 days ago
Comment by jmathai 3 days ago
But there are ways to control and constrain the LLMs and what the user is presented with.
These are all top of mind for me and why I felt there could be a better option than asking ChatGPT directly.
Comment by hirako2000 3 days ago
What's difficult and doesn't have to be with philosophy/ spirituality is to find relevant bits off situation, theme etc.
This app does that very well, LLMs are good at entity recognition.
Comment by weiliddat 3 days ago
Comment by swat535 3 days ago
I think it depends, Catholics wouldn't be able to use this because the Magisterium is the ultimate authority on interpreting Scripture, so the personal interpretation isn't really needed. This is not to say that Catholics don't read the Bible, they are encouraged to do so since it deepens their faith
On the other hand, for Protestant it varies, the High Church denominations are closer to Catholics (though none of them accept the Magisterium) in terms of scripture interpretation, but the Low Church ones (like Baptists or Non-Denominational ) are more open to personal interpenetration.
Disclaimer: I'm a Catholic, so if I made a mistake here fellow Protestants, please correct me.
Comment by foltik 3 days ago
Comment by ComputerPerson 3 days ago
I'm on a team that develops a Bible study app, and we're all relatively content with how the basic models converse regarding scripture. Even as far back as GPT-4 was excellent. They occasionally have minor hallucinations (a dealbreaker for a production app), but they do an excellent job with theology and Bible scholarship, given reasonable guardrails.
I'll admit I'm coming from the perspective of "should we be implementing this?" It seems, on the surface, that a strong embedding-based verse retrieval covers the bases at a microfraction of the cost.
If you're interested, check out the development server where we're working on this. You navigate to the search (magnifying glass) and then hit "Meaning". Sorry for the confusing route; we're still deciding on back-end details and haven't focused on the front yet.
Comment by willchis 3 days ago
Comment by swingboy 2 days ago
Comment by jmathai 3 days ago
So it’s less about model choice and more about governance of scripture.
I will check out the link you sent for sure!
Comment by ComputerPerson 3 days ago
Anyways, impressive app! We haven't tackled such an ambitious project just for it being daunting.
Comment by jmathai 2 days ago
Comment by irthomasthomas 3 days ago
Comment by jmathai 3 days ago
Comment by kfse 3 days ago
Comment by vmt-man 3 days ago
flash is suitable only for a toy apps, not for production environments :)
Comment by tmaly 2 days ago
Comment by jmathai 2 days ago
https://api-docs.deepseek.com/quick_start/agent_integrations...
Comment by achalxyz 2 days ago
Comment by websap 3 days ago
Deepseek v4 Pro prices with Opus 5 perf would be freaking unbelievable!!
This is probably a dream.
Comment by jug 1 day ago
Comment by lionkor 3 days ago
Comment by adrian_b 3 days ago
Comment by adrian_b 3 days ago
Comment by WithinReason 3 days ago
https://artificialanalysis.ai/models/deepseek-v4-flash?intel...
Comment by spwa4 3 days ago
First, your direct comparison, Deepseek V4 Flash 0731 (max effort) $0.03 (rounded up) per task @ index 50.
OpenAI Luna:
* high effort $0.03 (rounded down) @ index 46
* xhigh effort $0.04 @ index 49
* max effort $0.07 @ index 51
So I would say a fair statement would be "OpenAI Luna between 2x and 3x the price of Deepseek Flash, what you get is 2 to 5 times faster inference"
The cheapest OpenAI model that beats it is OpenAI Luna (max effort) $0.07 @ index 51 (if you take the rounding out it summarizes to triple the price for similar performance), but still close to 3x faster.
And can SOMEONE please tell artificialanalysis that using dark blue for both Deepseek AND OpenAI is an especially unfortunate choice of colors, especially today?
Comment by andai 3 days ago
For simple tasks, they're already saturated, and you'd prefer the faster model, so that you can have a realtime/interactive-ish experience.
Or to put it bluntly, it's cheaper if you don't value your time. That goes for smaller models in general -- need more handholding, more correcting -- but the Chinese ones are slower on top of that.
As for speed, Sol on Low is faster than Luna on most settings.
Comment by seaal 3 days ago
uhh openai is dark gray: `rgb(31, 31, 31)` and i'm pretty sure it always has been?
Comment by ComputerGuru 3 days ago
Comment by slopinthebag 3 days ago
Comment by ComputerGuru 2 days ago
Comment by soerxpso 1 day ago
Comment by akurilin 3 days ago
Comment by kamranjon 3 days ago
Comment by scosman 3 days ago
Plus a size you can genuinely run at home: Unsloth lossless Q8 at 162GB.
Comment by luckydata 3 days ago
Comment by cmrdporcupine 3 days ago
But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.
Comment by wolttam 2 days ago
Going local has as opened up a world of use-cases I never would have entertained the idea of on metered/cloud usage. Privacy is a large part of it but, I also no longer think twice about whether to send a prompt or not based on the psychology of it costing money.
Cached input tokens on local inference are free, so I don’t care about running sessions up to 500k tokens and hundreds of turns (it’s rarely useful, but DSv4 remains surprisingly coherent up there)
Comment by vardalab 2 days ago
Comment by ycui7 3 days ago
owning a few GPUs is a lot cheaper than supercars.
Comment by cmrdporcupine 3 days ago
I don't use it for local inference so much. I use it to learn.
I also use it as my daily driving Aarch64 development system.
Aside it's also very cool what else can be done with unified GPU memory, once you realize you have it...
Comment by bethekind 2 days ago
Comment by segmondy 3 days ago
Comment by scosman 3 days ago
Comment by sourcecodeplz 2 days ago
maybe it is Fable level
Comment by cmrdporcupine 3 days ago
If the full non-flash model follows up with the expected improvements, and at the price point they've been keeping, it puts the frontier labs in a tough position and it feels to me like like OpenAI is reaching deep into their pockets to try to head that off.
TFA link is a 404 though. I'm reading through https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731 instead
Comment by WhitneyLand 3 days ago
It’s also so inefficient, when they release the full performance numbers it’s not going to be good.
One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
Comment by Bnjoroge 3 days ago
Comment by drob518 3 days ago
Comment by kzrdude 3 days ago
Comment by WhitneyLand 3 days ago
Comment by prism56 2 days ago
Comment by WhitneyLand 2 days ago
All else being equal passing triple the amount of tokens through a model to solve the same problem makes it slower.
Doesn’t mean this model is bad, and it has to be considered how cheap it is, but it’s a factor.
Comment by rat9988 2 days ago
Comment by mdp2021 3 days ago
That depends. Is it also more reliable?
If two books, one big one slim, prove the same thesis, what I would be interested in is the quality of the content, not the size. There can be a measure of efficiency in "have you really thought it through", but it is clearly complex - it requires measuring how solid the reasoning is.
Comment by ptole_my 3 days ago
Comment by onlyrealcuzzo 3 days ago
Mind you, until the recent price cuts to Luna - Gemini 3.6 Flash wasn't even egregiously priced (but oh how things change in just 1 week).
Comment by ycui7 3 days ago
It use the stock model, no new models requires.
Worth spend a few hours to try.
The DGX Spark requires a small hack to ignore the difference between sm120 vs sm121, but it does run on sm121.
Comment by coder543 3 days ago
Comment by kamranjon 3 days ago
Comment by vmt-man 3 days ago
Comment by kamranjon 3 days ago
Comment by FuckButtons 3 days ago
Comment by hetspookjee 3 days ago
Comment by ycui7 3 days ago
Comment by tough 3 days ago
Comment by ycui7 3 days ago
Comment by theturtletalks 3 days ago
Comment by kamranjon 3 days ago
Comment by Lwerewolf 3 days ago
Comment by knuppar 3 days ago
Comment by walrus01 3 days ago
Comment by WithinReason 3 days ago
Comment by ttul 3 days ago
You learn something every day. Today, it was the term "systolic array": A systolic array is a specialized grid of simple, interconnected processing units designed to execute parallel data operations—like matrix multiplication—by rhythmically passing data directly from cell to neighboring cell without writing intermediate results back to main memory.
The term comes from the biological word systole (the contraction of the heart pumping blood through the body). In a systolic array, data "pulses" through a network of processing elements on every clock cycle, driven by a global clock beat.
Comment by pstuart 3 days ago
Comment by porridgeraisin 3 days ago
As for sibling comments, huawei ascend is more of an NPU-style architecture where you can easily have much bigger MMAs as primitive. But you usually don't anyways for many reasons.
Comment by andersa 3 days ago
Comment by ycui7 3 days ago
Comment by nolist_policy 3 days ago
Comment by Archit3ch 3 days ago
Comment by 1saadcodes 2 days ago
Comment by baalimago 3 days ago
The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.
Comment by Flere-Imsaho 3 days ago
Comment by ljosifov 3 days ago
Comment by jmartrican 3 days ago
Comment by wongarsu 3 days ago
Of course you can get more bang for your buck by being more deliberate. But that's equally true with US frontier models. You can optimize your work by choosing between Opus, Fable, Sonnet, Sol, Luna and Terra for each task. Some people seem to prefer to let Opus code and Sol review, for example. And then there is the whole debate whether $current_version is actually better (Some people stay on Claude 4.8 because they dislike how 5.0 is sometimes doing stupid things, just as many opted out of dynamic reasoning when they still could)
Comment by rented_mule 3 days ago
Comment by felixgallo 3 days ago
Comment by stavros 3 days ago
Claude is very spendy.
Comment by felixgallo 2 days ago
Comment by stavros 2 days ago
Comment by Flere-Imsaho 3 days ago
Comment by dominotw 3 days ago
Comment by 0xc133 3 days ago
Comment by Demiurge 2 days ago
Comment by dominotw 2 days ago
Comment by dominotw 2 days ago
Comment by ReptileMan 3 days ago
Comment by VulgarExigency 3 days ago
Comment by rapind 3 days ago
I've used it in some open source code though, and loved how fast it was.
My mind is changing on how valuable my code actually is though... it's the complete picture, how it's put together, the design, the UI, the attention to detail that's the real value.
Comment by nodja 3 days ago
Comment by refulgentis 3 days ago
Comment by singingtoday 3 days ago
I don't find it suitable for everything, but there's some tasks it crushes for what feels like almost free.
Comment by jingpostmedia 3 days ago
Comment by prathje 3 days ago
Comment by smrtinsert 3 days ago
Comment by f311a 3 days ago
Comment by esafak 3 days ago
The official release of the DeepSeek-V4-Flash API is now in public beta. The API calling method remains unchanged — simply set the model name to deepseek-v4-flash to use the latest version.
https://api-docs.deepseek.com/updates/#date-2026-07-31Comment by momojo 2 days ago
Comment by epolanski 3 days ago
The specific agent is focused on getting precise and on point answers about a codebase.
The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all.
The benchmark included more than 50 questions or different difficulty.
But when the agent was improved in its prompt and rooting it was impossible to have it perform worse than closed source sota.
Just to say that the quality of the harness is as important as agents intelligence.
Comment by quikoa 2 days ago
Comment by lostmsu 3 days ago
Comment by Catloafdev 3 days ago
Comment by simonw 2 days ago
But a REALLY good pelican on reasoning mode high (via OpenRouter): https://static.simonwillison.net/static/2026/deepseek-flash-...
Comment by apitman 3 days ago
Comment by mchusma 2 days ago
Comment by cheesecakegood 2 days ago
Comment by baq 2 days ago
Comment by cherryteastain 2 days ago
Comment by surgical_fire 1 day ago
DS has a lot more than mere 80% of Claude's capability, and the price is a lot less than 20%.
Comment by GaggiX 2 days ago
It's good.
Comment by baq 2 days ago
Comment by ignoramous 2 days ago
> who is the customer
If no one else, then definitely those that can only budget $1/mo to $5/mo. Probably 100s of millions, if not billions, of the Global South.
[0] Free access to DeepSeek v4 Flash is indeed available from providers like OpenRouter, OpenCode, and Freebuff (to name a few), but no ZDR.
Comment by baq 2 days ago
Comment by GaggiX 2 days ago
Comment by embedding-shape 3 days ago
Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point.
Comment by Lwerewolf 3 days ago
Comment by hxii 3 days ago
Comment by denismi 3 days ago
Comment by qtalen 3 days ago
Comment by nilsbunger 3 days ago
Comment by pornel 2 days ago
Comment by SomeHacker44 3 days ago
Similar price? Doesn't make sense. Maybe they meant power, capability or speed?
Comment by storus 3 days ago
Comment by fillskills 3 days ago
Comment by darknoon 3 days ago
Comment by comandillos 3 days ago
Comment by gorkemyildirim 3 days ago
Comment by sim04ful 3 days ago
Comment by paoliniluis 3 days ago
Comment by k1e 3 days ago
Comment by Computer0 3 days ago
Comment by net01 3 days ago
Comment by k__ 3 days ago
Why do the cache hit rates seem to vary so much between harnesses?
I use pi, which is very minimalist, and I get a hit rate of ~99%. Paying like $1 a day for Flash. Yet, the hit rate mentioned on OpenRouter is only ~79%.
Comment by drob518 3 days ago
BTW, this is one of the things that I really like about Pi. It’s very simple and thus very predictable.
Comment by guess__who 2 days ago
Comment by shostack 2 days ago
What am I missing?
Comment by tmikaeld 3 days ago
[0] https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
Comment by pornel 2 days ago
Comment by sschueller 3 days ago
Comment by hnsmomdpvp 3 days ago
Comment by mrnobody_ 3 days ago
Comment by buildinext 3 days ago
Comment by Aeroi 3 days ago
create a plan with SOTA, execute with this.
Comment by NooneAtAll3 3 days ago
Comment by bigmadshoe 3 days ago
Comment by WithinReason 3 days ago
Comment by theanonymousone 3 days ago
Or a benchmark to benchmark benchmarks?
Comment by NooneAtAll3 3 days ago
benchmark website benchmark is indeed a benchmark that benchmarks websites with benchmarks (but it can be shown outside websites as well, it's not picky)
Comment by sourcecodeplz 2 days ago
Comment by 0xchamin 3 days ago
Comment by hjm 2 days ago
...touching myself RN.
Comment by try-working 3 days ago
Comment by _ache_ 3 days ago
They can't compete, they have bills. By the end of the year, if they can't react, it's game over.
Maybe US clients could be a little patriotic here, but money is money. They won't give them free money forever.
Comment by try-working 2 days ago
Comment by freakynit 3 days ago
The ban on these open models is coming within weeks, if not days. As usual, the excuse will be "national security".
Comment by Der_Einzige 3 days ago
I claim the CCP will wise up within 2 years, possibly much much sooner, and ban their own companies from open sourcing to prevent the Americans from acquiring the capabilities.
Despite all the nonsense claims of China distilling US models, the reality is that the Americans absolutely do distill these free Chinese models, and distillation when full logprobs are available (i.e. you have access to the weights of the model) is an order of magnitude better than when you don't.
Yes, Chinese open weight models in the short term harm US closed source model providers bottom line. In the slightly longer term, "showing your hand" and publishing both the architecture innovations and the models weights will be too dangerous for the CCP to allow. This is triply true if they can release a model that beats the Americans on most benchmarks.
I've already warned investors that this is probably the closest open weight models will ever get to closed access.
Comment by VulgarExigency 3 days ago
https://www.businessinsider.com/xi-jinping-open-source-ai-us...
Comment by mcbuilder 3 days ago
Comment by rapatel0 3 days ago
Comment by UltraSane 3 days ago
Comment by freakynit 3 days ago
For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models.
They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for such open models.
Comment by qphe95 3 days ago
Comment by nancyminusone 3 days ago
Comment by tyfon 3 days ago
I'm not sure the outcome would be beneficial for the US as a whole here. But perhaps that is not their priority.
Comment by freakynit 3 days ago
Comment by UltraSane 3 days ago
Comment by freakynit 3 days ago
Comment by dgellow 3 days ago
Comment by UltraSane 3 days ago
Comment by dgellow 3 days ago
Comment by UltraSane 3 days ago
Comment by hgoel 3 days ago
Comment by UltraSane 3 days ago
Comment by dgellow 2 days ago
- the US strictly regulate cryptography https://en.wikipedia.org/wiki/Export_of_cryptography_from_th...
- some prime numbers are considered illegal https://en.wikipedia.org/wiki/Illegal_number#Illegal_primes
I don’t think banning open weight is that far fetch compared to those
Comment by UltraSane 2 days ago
Comment by dgellow 2 days ago
Comment by dgellow 3 days ago
Comment by Der_Einzige 3 days ago
People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!
Comment by nickthegreek 3 days ago
Comment by freakynit 3 days ago
Since when have hackernews started to become toxic like stackoverflow used to be?
Comment by dgellow 3 days ago
commenting about voting is also something the HN guidelines warns against:
> Please don't comment about the voting on comments. It never does any good, and it makes boring reading.
Comment by bellowsgulch 3 days ago
Comment by monooso 3 days ago
Comment by theanonymousone 3 days ago
Comment by theanonymousone 3 days ago
Comment by brynnbee 3 days ago
Comment by mlmonkey 3 days ago
Comment by WithinReason 3 days ago
Comment by ValentineC 3 days ago
Comment by madikz 3 days ago
Comment by BedVibe_Studios 3 days ago
Comment by yucongchen 3 days ago
Comment by iluvcommunism 3 days ago
Comment by dpacmittal 3 days ago
Comment by dang 1 day ago
Comment by hilios 3 days ago
Comment by seanmcdirmid 3 days ago
Comment by shwetanshu21 3 days ago
All the good talent moved to other countries due to this.
Comment by budsniffer952 3 days ago
Comment by procgen 2 days ago
Comment by seanmcdirmid 2 days ago
Comment by otabdeveloper4 3 days ago
Comment by Wherecombinator 3 days ago
Comment by ndkddmmfmfm 3 days ago
Comment by LoganDark 3 days ago
Comment by baalimago 3 days ago
Comment by chronogram 3 days ago
Comment by speedgoose 3 days ago
Comment by seanmcdirmid 3 days ago
Comment by serial_dev 3 days ago
Comment by Gud 3 days ago
Comment by speedgoose 3 days ago
Comment by tao_oat 3 days ago
Comment by ungovernableCat 2 days ago
But if models will be commodities, maybe that won't be such a terrible thing longterm?
Comment by Barrin92 2 days ago
The fact that China could catch up in what, two years, shows that this is a commodity.
Comment by lerchmo 2 days ago
Comment by m00dy 3 days ago
Comment by Nifty3929 2 days ago
It needn't be competitive - everybody can succeed more tomorrow than they are today. You win. I win. We win.
Comment by Saline9515 3 days ago
Comment by InsideOutSanta 3 days ago
Comment by codedokode 3 days ago
But imagine if Chinese labs would lose. Then there will be no open models, and the prices for closed models would be raised to the maximum.
Comment by shimman 2 days ago
Comment by donquichotte 3 days ago
Comment by throwa356262 3 days ago
Comment by donquichotte 2 days ago
Retaining world-class talent and home-growing industries is no small feat. In Europe, we have been struggling with this, and I am not even entirely sure why.
Comment by dannyw 2 days ago
Comment by singingtoday 3 days ago
Comment by slopinthebag 3 days ago
Comment by segmondy 3 days ago
Comment by dominotw 3 days ago
Comment by dpacmittal 3 days ago
Comment by dominotw 3 days ago
yea i got that from your first comment ( although you removed crush American companies in _price_ ). you are pro cheapness at any cost even if its from your country's state funded direct geopolitical enemy.
China can always count on first order greed to win
Comment by satvikpendem 3 days ago
Comment by VulgarExigency 3 days ago
Comment by dominotw 2 days ago
Comment by VulgarExigency 2 days ago
Comment by dominotw 3 hours ago
Comment by ungovernableCat 2 days ago
Comment by xXSLAYERXx 3 days ago
Comment by johnny_rico 2 days ago
Sanctions and trade policies that hurt ordinary people? Cultural dominance through Hollywood (Holy Wood, again with the sick ruling elite and their witchcraft crap), social media, brands, the poisoning of food and water on global scale?
Etc? Etc?
Most people don't hate "Americans" or their companies, they are just against your ruling elite. Live one day outside the united states and understand how the planet works for the rest of mankind and you might start to grasp on reality.
And yeah, China will catch up with hardware pretty damn soon. Jesus, you guys bought your radios from Japan not 50 years ago. Sorry, can't dominate the world's economy forever. Never happened in history.
No matter how much black magic, dark rituals, sick crap and sacrifices people on the top throws at the wall.
We're humans, after all.
So yeah, it is not against America itself or American companies, but against the ruling class. So naturally , people get excited about China entering into the match!
3...2...1... Fight!
Comment by xXSLAYERXx 2 days ago
Most everyone wants to come to America for the opportunity. Think an immigrant in China has the same economic opportunity as in America? Any other country come to mind?
Comment by eunos 2 days ago
Comment by xXSLAYERXx 2 days ago
Comment by antonvs 2 days ago
Comment by xXSLAYERXx 2 days ago
Sure, if you spend most your time on the internet in various news forums its doom and gloom. Its what sells.
Comment by Danox 2 days ago
Comment by xXSLAYERXx 2 days ago
Comment by AmazingTurtle 3 days ago
Comment by Der_Einzige 3 days ago
Daily reminder that improving your samplers from the garbage default top_p/top_k to min_p or subsequent methods dramatically improves the performance of these models, and makes most quantities like measured "verbosity" and subsequent calculations of "intelligence per token" meaningless
Daily reminder that no one, including within academic AI research, AI engineers, normies, etc takes LLM sampling seriously enough.
Comment by JSR_FDED 3 days ago
Comment by bawana 3 days ago
Comment by yonisto 3 days ago
Comment by xbmcuser 3 days ago
Comment by edot 3 days ago
Comment by nancyminusone 3 days ago
I have to admit it rarely comes up in the coding tasks I usually give to LLMs.
Comment by Boxxed 3 days ago
Comment by nancyminusone 3 days ago
Comment by avazhi 3 days ago
But you already know that.
Comment by SJMG 3 days ago
Comment by zawaideh 3 days ago
Comment by cogman10 3 days ago
If you poke it just a few times, however, you get to the point where it will eventually say (paraphrasing) that basically only Israel, the US state department, and the ICJ say it's not a genocide.
That is to say that it's framing it as some sort of tricky complex question when it's not. And when interrogated, it basically admits that the only people who dispute it are Israel and it's supporters.
Comment by throw59525773 3 days ago
The majority of the world is religious - doesn’t mean the debate on religion isn’t a complex question.
The majority of the world approved of slavery historically.
The majority of countries have ethnically cleansed their Jews, many of them in living memory.
When interrogated you will find that the only ones asserting the war in Gaza is a genocide are people who were anti-Israel anyway.
Comment by cogman10 3 days ago
I'm always suspicious that's the case given how mentions of gaza seem to bring out brand new accounts who only talk about Israel.
[1] https://quincyinst.org/research/the-eighth-front-inside-isra...
Comment by throw59525773 3 days ago
We are on a thread discussing Chinese models. Every discussion on here that’s negative about China or its models suddenly gets derailed via whataboutism to Israel/Gaza. A very convenient distraction.
Comment by cogman10 3 days ago
And yes, we were discussing censorship of models which, as I pointed out, doesn't seem like ChatGPT is directly censoring data though it does appear to be manipulating it. Pretty on topic.
It was you, brand new account hiding your past opinions, who came in here to make this solely about Israel.
Yeah, I think you are likely a foreign agent. Prove me wrong and post from an established account.
Comment by avazhi 2 days ago
Under no definition of genocide (as defined either going back to the middle 20th century or as that term is defined legally today) is Israel committing genocide. Simply put, there is no evidence that Israel is committing "specific acts with the intent to destroy, in whole or in part, a national, ethnical, racial, or religious group." Intent is, as with almost everything in the criminal law, everything here. If you want to subdivide Muslimes in a way that results in, say, Hamas and Hezbollah being their own religious groups then you could mount the argument on that basis, but those are political designations and so good luck with that. And of course neither of those groups encompass the entirety of the national polities they claim to represent - just look at how hated Hezbollah is in Lebanon. They are being ostracised and even the Lebanese leaders want nothing to do with them at this point.
People like you degrade actual acts of genocide such as what happened with the Nazis or such as what has happened several times in SE Asia and Africa over the past 50 years. A country acting in self-defense is not genocide no matter how much it might hurt your sensibilities. Genocide is, as is obvious to most people, extremely difficult to prove. See Armenia for how fraught this debate is. But acts of brutality in a conflict or an extreme imbalance of military power between belligerents alone do not mean that genocide is occurring.
Good day.
Comment by tarasn 3 days ago
Comment by shalom1112 3 days ago
Comment by vehemenz 3 days ago
Comment by adrian_b 3 days ago
Any normal user is much more likely to ask questions to which the Anthropic and OpenAI models do not answer, than to ask questions about the modern Chinese history, to which a Chinese LLM will not answer.
Comment by rescbr 3 days ago
This has been debunked here on HN so many times. The Chinese open models do answer the hairy Chinese political questions, and the raw APIs pass-through the response. Now, the answer might be blocked by the agent who's calling the API, specially if you are using a Chinese endpoint instead of the RoW (i.e. Singapore) endpoint.
That's the reason why you should always prefer a open agent/harness as well instead of using the provider's.
Comment by krull10 2 days ago
Comment by net01 3 days ago
Comment by DubzCheckEm 3 days ago
Comment by hiherer 3 days ago