Kimi Linear: An Expressive, Efficient Attention Architecture (2025)
Posted by ronfriedhaber 6 days ago
Comments
Comment by senko 6 days ago
Comment by bratao 6 days ago
Comment by iandanforth 6 days ago
Comment by muricula 5 days ago
Comment by throwa356262 5 days ago
The nvidia version is heavy to compute (common problem with RNN and LSTM, also noted in the paper). Moonshot's Kimi K3 replaces part of it with some function that performs better.
I dont understand the details but, that in itself might a pretty big contribution.
Anyway, I'm amazed how fast these companies improve each other's ideas and put them in new products.
Comment by pooyamo 6 days ago
Like isn't it weird that the 1 million parameter model with the same architecture can't solve basic puzzles but suddenly the 1 trillion parameter can conjure up counter-examples for the Jacobian conjecture?
It's unintuitive since, to the best of my knowledge, one of the basic tenants of algorithm development was that you can't just brute-force your way towards a solution for some complex problems, e.g. naive sorting algorithms suddenly won't beat quicksort if you put more processing to them, but in the modern LLM scene it seems people are in a race to scaling up, experimenting empirically and hoping the same algorithm/architecture comes to a solution.
Comment by jlamberts 6 days ago
> One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning.
The full essay is worth a read, it's pretty short http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Comment by BinaryRage 5 days ago
Comment by verdverm 6 days ago
Comment by pachev 6 days ago
Comment by entrepy123 5 days ago
Comment by realitylabs 5 days ago
Comment by verdverm 5 days ago
Comment by boilerupnc 5 days ago
Comment by IsTom 6 days ago
Comment by verdverm 6 days ago
Comment by zparky 6 days ago
Comment by verdverm 6 days ago
Comment by zparky 5 days ago
Comment by verdverm 5 days ago
Comment by cubefox 5 days ago
Comment by verdverm 5 days ago
Comment by gct 5 days ago
Comment by ilikehurdles 4 days ago
Comment by hnfong 6 days ago
It's kind of sad that popular CS textbooks often focus on solving precise problems with lowest theoretical complexity bounds while ignoring more practical (but generally applicable) computation techniques.
In machine learning they call it "gradient descent", which in older days had analogies in techniques called "hill climbing", "local search" and "simulated annealing". Basically you have a function you need to optimize for, and you clumsily tweak the parameters so that you get the (locally) max/min value you wanted. These techniques were great at finding approximate, locally maximal solutions without trying all the possibilities at once (which is more akin to the kind of "brute force" in the traditional CS context).
I guess because these techniques were generally applicable yet the outputs were approximate and you couldn't analyze them much (no fancy O(n log n)), the theorists did not find them interesting and thus were not put into the spotlight of student's learning curricula.
In modern machine learning they do this gradient descent thing which is also tweaking the parameters bit by bit to optimize for the loss function, except that the parameters are now in the billions and trillions. The compute required is huge of course, but it's actually quite an "efficient" process, and it's not actually doing much of "brute forcing" at all. During training, the process is essentially, almost equivalent to, compressing the many many trillions of tokens of training data. To me it's quite amazing that they manage to complete such a process within a couple months of training, even if they have hundreds of thousands of GPUs...
Comment by nuancebydefault 5 days ago
I thought this is nowadays called gradient descent
Comment by IanCal 6 days ago
Going from 1m to 1T params is also a scaling of a million times. It’s like going from a human brain down to one percent in size in each direction or just a few mm.
Comment by joefourier 6 days ago
I'm not sure what you mean? You can see the intelligence of LLMs progress predictably and stably according to scaling laws. LLMs have to encode language in addition to intelligence so there's a minimum bound for them to output sensible text (you can train specialised tiny models to solve basic puzzles without language). Start at around 127M and compare models of increasing parameters and you'll see a clear progression in intelligence.
> It's unintuitive since, to the best of my knowledge, one of the basic tenants of algorithm development was that you can't just brute-force your way towards a solution for some complex problems, e.g. naive sorting algorithms suddenly won't beat quicksort if you put more processing to them
How is that a basic tenet? Simple, easier to parallelise algorithms that have lower memory requirements, or can take better advantage of hardware, or don't hit a plateau the more compute you throw at them, can absolutely beat cleverer algorithms. E.g. brute forcing rendering with Monte Carlo path tracing will give you more physically accurate results than ray tracing or rasterisation algorithms that rely on a bundle of hacks to approximate global illumination, transparency smooth shading, etc.
Comment by sota_pop 5 days ago
It is _intelligent_ in the sense that it optimizes a thing that is hard for humans to not anthropomorphize.
Things like percolation theory, swarm theory, and others are similar topics in “complex systems”. Neural networks are interesting because they combine aspects of both complex systems and dynamical/adaptive systems (ie systems with a feedback loop).
A neural network at its very core is a function fitting algorithm. It stores matrices of parameters (ie weights and biases) such that every parameter (and combinations thereof) captures a relationship of your data in exactly the same way the slope and intercept are obtained through Linear Interpolation. All of the Regularization tricks applied just are attempts to incentivize a given parameter to not encode any trends too specifically (think an instance vs a type).
In this way you can ask yourself “how many aspects of your data are required to capture it adequately?” This is what scaling offers.
Comment by pornel 6 days ago
https://en.wikipedia.org/wiki/Grokking_(machine_learning)
High-dimensional gradient descent behaves very differently than the simplified 3d visualisations we use to demonstrate it, and has lots of ways out of local minima:
https://www.youtube.com/watch?v=NrO20Jb-hy0
so it seems like there is a benefit to giving models more space to learn in rather than forcing them to compress the knowledge from the start.
Comment by CamperBob2 5 days ago
Comment by neerajk 5 days ago
And as scaling reduces a model's excess entropy, the model can become good at combinations of skills much faster than you would expect if it had to separately see and memorize every combination. They call this "slingshot generalization".
Comment by coderenegade 5 days ago
The bigger model has a better chance of gaining a foothold in that internal representation space where inputs are mapped to meanings and outputs, and eventually, it optimizes to the point where most of the weights aren't doing much. It's not immediately clear how much expressivity is required by the network to learn that space, but so far the answer seems to be in the billions of parameters.
The more interesting question to me is to what extent we should expect the models to be invariant to data. For instance, if I learn a certain type of analysis, that skill shouldn't depend on the data I'm looking at-- it should be repeatable for any given data of the same type/class. I'm curious to what extent skills are embedded in the weights versus data and "facts". My hunch for why mathematical reasoning and programming resulted in large step changes in model performance across the board is because these are inherently skills that are widely repeatable for a large class of tasks. The ability to express programmatic logic is invariant to both the language and the task at hand. And to me, that's how you get to smaller models: by focusing on the skills.
Comment by hodgehog11 5 days ago
To be clear, it is an extremely narrow model class that can do this; we just got "lucky" and worked our way to it. That's why we still teach general statistical principles which often forbid this sort of behavior as a rule of thumb.
Comment by thomasahle 6 days ago
Let's say there's some circuit that does problem solving of the kind we call intelligence.
We dont know what this circuit looks like, but it exists in our brain.
Doing regression on outputs from the brain (e.g. internet text) with enough parameters, we can "fit" our model to this circuit.
But if you try to fit it with fewer parameters than it needs, you're just going to get some linear approximation.
Comment by hendiatris 6 days ago
Comment by esafak 6 days ago
Comment by lacunary 6 days ago
Comment by lern_too_spel 6 days ago
Comment by chpatrick 6 days ago
Comment by ACCount37 5 days ago
In a way, training an AI is: using an algorithm to find, discover and refine other algorithms computationally. When we scale, we pour more raw inputs into finding the right algorithms for a given objective. Is it surprising, then, that we find a better algorithm for it?
As a very stupid analogy - for a given task, a small model's internal algorithm might, under its capacity and training signal constraints, top out at slightly above "bubble sort". While a larger one could dig deeper and get closer to "quicksort" internally. A naive, dirty algorithm got replaced by a more sophisticated algorithm that performs better.
Keep in mind: intelligence is very much not a binary. There's no "threshold" at which a model goes from "this is dumb statistics" to "this is actual intelligence". A 1B LLM and a 10T LLM both have some amount of intelligence. It's just that one would have intelligence that's so weak, underdeveloped, and overindexed on statistical regularities that it's very easy to dismiss it altogether. And the other might have enough of it to snipe unresolved conjectures with novel counterexamples in math. Makes it considerably harder to dismiss it outright. While the curve between the still two looks less like an abrupt jump and more like little increments building up to an avalanche.
The jump in something advanced and specific like "math abilities" can look quite sharp - but under it, there are far more generic capabilities that back it. They build up slowly to eventually enable that jump. A more advanced model makes less reasoning mistakes and recovers from reasoning mistakes more gracefully - two very generic capabilities - but once those capabilities improve enough, a whole new type of logic problem might fall to it.
Comment by storus 5 days ago
Comment by 3836293648 4 days ago
Comment by mohsen1 6 days ago
Mote-Carlo is pretty useful still. Not sure if your statement holds
Comment by knaeckeKami 5 days ago
Yes, you can't compare biological neurons to parameters in an LLM 1:1, neurons do much more, but still - It is plausible that higher intelligence just requires scale.
At least not even evolution over millions of years seems to have found a way to produce human-level intelligence in a smaller scale, even though it had a lot of incentives to do so (the human brain needs a lot of kcal).
Comment by Mr_Minderbinder 5 days ago
Comment by lacunary 6 days ago
Comment by berz01 6 days ago
Comment by oakpond 6 days ago
This is just awesome.
Comment by Cort3z 5 days ago
Comment by trollbridge 5 days ago
Comment by delichon 6 days ago
Comment by egeozcan 6 days ago
Comment by reilly3000 6 days ago
Comment by idiotsecant 6 days ago
Comment by petu 6 days ago
https://xcancel.com/teortaxesTex/status/2026130112685416881
I think I've seen same happening with some European languages as well.
Comment by snovv_crash 5 days ago
Comment by __MatrixMan__ 6 days ago
Comment by overfeed 5 days ago
Comment by halJordan 5 days ago
Comment by egeozcan 5 days ago
Distillers are not taking anything, they are just making their model learn from better ones - isn't that the whole AI training doesn't violate IP argument?
They just aren't using the tool in compliance with the terms of service. Anthropic could ban them or take them to court maybe. Not an attack still.
Comment by Tostino 5 days ago
Comment by anigbrowl 5 days ago
...or maybe we stop defaulting to adversarial paradigms for every conceivable situation.
Comment by nextaccountic 5 days ago
A model is teaching another here
Everyone wins
Comment by WhitneyLand 6 days ago
Are Chinese labs impressively innovating? Clearly.
However this doesn’t rule out possible gains due to distillation.
I don’t know the degree of the latter but both things could certainly be true.
Comment by SirHackalot 6 days ago
Comment by DashAnimal 6 days ago
Comment by SirHackalot 6 days ago
Comment by hugopuybareau 6 days ago
Comment by joe_the_user 5 days ago
The process of creating an LLM involves taking and processing a massive amount of human generated data, roughly all the world's literature/thinking/etc. A large portion remains within these systems. Aside from the legality, ethically that shouldn't belong to any one company.
Comment by SirHackalot 6 days ago
Comment by linkregister 5 days ago
Comment by hoihoi 5 days ago
Comment by joe_the_user 5 days ago
Comment by fnord123 6 days ago
Comment by trollbridge 5 days ago
Comment by logicallee 5 days ago
Comment by fnord123 5 days ago
> The requirements set forth in Sections 2 and 3 do not apply to: (a) internal use of the Software, defined as any use that does not make the Software, its outputs, or its underlying capabilities available to third parties; or (b) any use of the Software accessed through Moonshot AI's official products or certified inference partners.
"distilling" from their own hardware would be an internal use of the software.
Comment by logicallee 5 days ago
Comment by cma 6 days ago
Comment by culi 6 days ago
Comment by api 6 days ago
“You stole my warez!”
Comment by serial_dev 6 days ago
Comment by esafak 6 days ago
Comment by Barbing 5 days ago
Comment by trollbridge 5 days ago
Comment by igleria 6 days ago
Comment by Parfait__ 6 days ago
Comment by fwip 6 days ago
Basically, they want IP protection for Claude. This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.
Comment by blints 6 days ago
Comment by verdverm 6 days ago
https://docs.cloud.google.com/gemini-enterprise-agent-platfo...
Comment by petu 6 days ago
It's "optimize your costs in our garden" product.
Comment by verdverm 5 days ago
I recommend Fireworks as an alternative
Comment by behnamoh 6 days ago
Comment by verdverm 6 days ago
Comment by charcircuit 5 days ago
Comment by koe123 6 days ago
Comment by fwip 6 days ago
Realer answer: A combination of the above, plus group/bubble effect of all your coworkers saying the same thing. You as a group, conflate a bunch of concerns together (China, no-guardrails-AI, etc), decide that your group will be the responsible stewards of AI, and then anybody "stealing your work" appears dangerous - both to your livelihood and to the human race.
Comment by Levitz 6 days ago
No, it's perfectly reasonable once you get down to reality.
China is not going to care about IP. That's just a fact. So either nobody cares about IP (at the very last in this context) and any AI company can just do whatever with data, or Chinese companies have to be held up to scrutiny.
We don't have the privilege to be able to hold western companies to higher ethical, legal, and environmental standards and not risk competitiveness.
That there is a whole lot of people right now who insist on doing above and still somehow praise China at every turn is something historians or news outlets will have to make sense of in some 5 years time.
Comment by nemomarx 6 days ago
Comment by fwip 5 days ago
Anthropic is currently angling for a "IP restrictions for China but not for me" la-la land scenario. If they were instead arguing for either of the choices you said, it wouldn't be hypocritical.
Comment by idiotsecant 6 days ago
Comment by creato 5 days ago
So distillation, among other things, allows Chinese labs to indirectly benefit from these arrangements at no cost to them.
Comment by trollbridge 5 days ago
Comment by lukewarm707 6 days ago
the chinese models are less censored, you'd better believe it.
try asking claude about its 'guardrails' (restrictions), very high chance anthropic will censor it.
Comment by Barbing 5 days ago
Tried spicy geopolitical dispute questions?
Comment by trollbridge 5 days ago
Comment by Barbing 5 days ago
This user tried Qwen (and DeepSeek, Kimi, GLM) ten days ago, called the response's language milquetoast: https://news.ycombinator.com/item?id=48964345
Comment from a user a month ago with their blog link on DeepSeek, which for comparison includes a generated poem about the evil killings of unarmed students by the USA at Kent State: https://news.ycombinator.com/item?id=48680612
(I vouch for none of these alleged tests.)
Of course when these tests used hosted models, different rules can apply. btw China is an amazing country with amazing people; my interest is in a little of everything, including free SotA-adjacent open software (thank you China!). Both respect their sovereignty to regulate software as used within their borders and hope labs there find value in serving Western users with the openness/transparency we (I) like to think we desire. I believe grappling with the most uncomfortable topics will be to their advantage in the long run. And ours (USA), too. The more we can divorce ourselves from bias, the more we can all win, I hope. My bias is towards humanity [being safe, happy, fulfilled...].
(While I might sound like someone who'd appreciate the model billed as "maximally truth seeking", unfortunately due to training or system prompting or something, Grok is trash unless you need to search Twitter or perhaps bypass botblocks. Or make "7K sex images of stepdaughter", I reference with apologies and deep sympathy to Jane Doe 4.)
Comment by lukewarm707 5 days ago
https://huggingface.co/spaces/DontPlanToEnd/UGI-Leaderboard
the stock cn models do have refusals, although fewer than us models. critically the cn models are mostly open weights, many in the top 10 UGI are uncensored models based on open weights. the weights and the information locked up in them are out there. abliterated google gemma does very well.
imagining the stock blank system prompt gives activations defensive of tibet; if you prompt it "you are john bolton" you will activate the opposite.
with gemini, openai, claude, you simply cannot, they will always be restricted and will restrict you from access to ai. they will restrict you from having access to the raw ai, which you can edit, tune, and apply to your needs.
the cn models are private because anyone can host them in any jurisdiction, i run all of them with zero data retention.
by contrast the moment you sign up for chatgpt openai and anthropic are spying on you, reading and storing your prompts and sending them to moderators when you violate their policies, even reporting people to the police.
as a side note, in regard to my own politics, i believe in life, liberty and the right to bear arms. i can't believe we are infantilising users with safety guards, spying on users for policy violations and trying to restrict open intelligence.
Comment by Barbing 4 days ago
Comment by HDBaseT 5 days ago
Comment by Aurornis 6 days ago
Comment by EGreg 6 days ago
https://www.bbc.com/news/technology-12343597
Microsoft replied that Bing uses “many different signals” —- including cribbing from Google :-)
Comment by krlx 6 days ago
Comment by moralestapia 6 days ago
The distillation theory does not even make sense as Fable was only around for days (effectively) before Kimi was released.
Comment by jeremyjh 6 days ago
Comment by vikramkr 6 days ago
Comment by rdtsc 6 days ago
Comment by imrozim 6 days ago
Comment by thesiti92 5 days ago
Comment by Topology1 6 days ago
Comment by KHUSHIL_GAMER 5 days ago
Comment by jasonjmcghee 6 days ago
As it's 9 months old and they just had a major model release
Comment by throwa356262 6 days ago
The main contribution of the K3 paper is Stable LatentMoE. Like some other models it compresses data sent between layers, which puts certain requirements on the router. K3 improves performance by using a more balanced expert selection strategy.
Comment by mcbuilder 6 days ago
Comment by throwa356262 6 days ago
Comment by senko 6 days ago
Comment by verdverm 6 days ago
the pretraining (slurping the internet) only goes so far, the new data being used is from human preferences and agent traces (designed and/or distilled)
Comment by cptcobalt 6 days ago
Comment by GaggiX 6 days ago
Comment by yorwba 6 days ago
Comment by GaggiX 6 days ago
Comment by yorwba 5 days ago