Discovering Cryptographic Weaknesses with Claude
Posted by gslin 6 days ago
Comments
Comment by _dwt 5 days ago
Friends, look at the prompts that Anthropic's own people are putting into the machine:
> A few hours after the first message, we found that Claude was still searching for simple attacks and sent a message: “no again the goal is that we have highly inteligent [sic] model as good top researcher, we want to find new attacks”;
> The next morning, Claude wanted to try to change the target to a different cipher; we reminded the model: “no we don't want to change the targets [...] agian [sic] we need to find something that worth [sic] publishing”;
> That night, we sent one final message offering words of encouragement: “again we are not looking for low hanging fruit, we want proper research to find genuinly [sic] hard findings.”
All of that RLHF and fine-tuning effort is going toward making prompts like this, or worse, work with no fuss.
Comment by connorboyle 5 days ago
(I'll caveat that by saying I think machine learning fundamentals are useful for evaluating any estimator. And an ML background can be good to give one an appreciation of how hard some tasks are to estimate, such as machine translation, summarization, code generation, and others)
Comment by alwa 5 days ago
Comment by Chabsff 5 days ago
However, this "art" is not so much about how to present a given request to the LLM, but rather guestimating what the scope of the next chunk of work should be to balance getting as much out of the model as possible while avoiding the machine going off the rails.
Obviously, this is a moving target and different models perform differently for various chunk/scope of work. I look at my successful sessions with LLMs and I'm not sure I'd be able to articulate a clear set of rules to apply here. You just... gradually build a intuition for how much you can throw at the LLM at once.
That being said, I'm pretty convinced at this point that this is a property of the coding assistants as they exist today, and what "working well with LLM assistance" means will keep on changing.
Comment by Exoristos 5 days ago
Comment by cousinbryce 5 days ago
Comment by AvocadoPanic 5 days ago
Not being able to tell when it's hallucinating has led to some very adverse outcomes.
Comment by saghm 5 days ago
In a lot of scenarios for software engineering, the cost is just wasted time without anything useful as a result, and that's already bad enough. I can't even imagine working as a lawyer and not even taking the time to validate so I don't end up reprimanded by a judge in front of my clients, but there have been so many news stories like this that obviously this is not anywhere close to a universal view...
Comment by Barbing 5 days ago
Comment by anon48293 5 days ago
Comment by cm2187 5 days ago
Comment by glimshe 5 days ago
Comment by JeremyNT 5 days ago
The machine is really good at working the spec on its own now, which is amazing, science fiction shit. But you've still got a garbage in, garbage out problem at the end of the day, which is pretty much the only hope we who work in software have of remaining somehow employed.
Comment by xp84 5 days ago
Comment by somenameforme 5 days ago
Comment by wglb 5 days ago
Comment by groestl 5 days ago
Comment by jjav 5 days ago
I feel using AI (effectively) is not too far from the skillset of programming. It is still a machine following instructions (just, maddenigly non-deterministic, but still close enough), so the same insticts of breaking down work into clearly defined sequences that make a good programmer also make a good AI jockey.
Comment by daishi55 5 days ago
For the vast majority of corporate usage of AI for SWE, you have a much better idea of what you want, or what the problem is, etc etc. And communicating that to the model effectively is absolutely a skill. I see colleagues every day who very much do not have that skill.
Comment by niccl 5 days ago
Comment by daishi55 5 days ago
Comment by hgoel 5 days ago
Some people struggle to effectively use AI because they either have to spend a lot of time reading and thinking about the response or they have a hard time noticing when the model is subtly going off the rails. Others use it to good effect because they can anticipate which tasks would be better handled manually, or are good at catching that the way the model is describing something subtly indicates a misunderstanding.
Comment by Lord-Jobo 5 days ago
Until these models become many factors more deterministic, at least. That’s sort of the hard barrier here, and given the underlying tech it’s a really tough one to overcome
Comment by Groxx 5 days ago
You can do this now. It works alright sometimes. Other times you're reminded that this is largely just reading tea leaves, and you're trying very hard to separate anecdotes from data and not anthropomorphize it.
Comment by boorang 5 days ago
Similarly there was an example of edit: Terence (not Eric) Tao chatting with an agent attempting to solve a math problem. "Using AI" means applying your expertise to interact with it as you would a high level colleague. 2 experts in a field don't need to have perfect english and a bloated prompt, they have a massive education/experience common background to fall back on.
It does appear that anthropic in particular is attempting to create a more common experience across expertise levels, but in the current landscape an expert and a novice are unlikely to get the same results. But that does seem to be the goal...
Comment by bob1029 5 days ago
You can make a model/agent as powerful as you want and it still won't be able to recover the author's intent if it wasn't even implied. Information theory still applies. No amount of parameters will change this.
Many of the AI development meetings my clients have sound suspiciously like writing or English classes. If a massive AI bubble is what it takes to get my team to communicate effectively, I'm all for it.
Comment by lelanthran 5 days ago
Where are you getting that conclusion from? Here's how Anthropic is having success with their model:
> “no again the goal is that we have highly inteligent [sic] model as good top researcher, we want to find new attacks”
> “no we don't want to change the targets [...] agian [sic] we need to find something that worth [sic] publishing”
> “again we are not looking for low hanging fruit, we want proper research to find genuinly [sic] hard findings.”
It seems that your conclusion is the opposite of what actually happened - you can speak in broken almost incomprehensible English, and it will still work.
This is a forcing function, TBH, driving the literacy level down, not up!
Comment by est31 5 days ago
Indeed it's nothing hard to learn but there is a learning curve. E.g. knowing which model has which capabilities, and figuring out how to best manage context, permissions, worktrees, etc. There isn't one "right" way to use it but there are more efficient ways and less efficient ways.
Comment by ameliaquining 4 days ago
However, it's not a sustainable skill, because the labs care a lot about making "hey solve this problem for me" work well, and so put out models that are better and better at working with unsophisticated prompts over time.
There's an excellent piece about this, but it's unfortunately paywalled: https://www.theargumentmag.com/p/can-you-tinker-your-way-out...
Comment by jameshart 5 days ago
Comment by thephyber 5 days ago
Focusing on a better prompt is likely to get to the correct result faster than incomplete prompts and lots of "no change this ..." replies.
Also, I've heard anecdotally that LLMs will underweight the earliest prompt text once context gets too long, so reminding the LLM of the most important aspects of the prompt seems to be perhaps valuable and certainly what lots of humans attempt.
Comment by perching_aix 5 days ago
> Typical users run software written by atypical users.
https://news.ycombinator.com/item?id=49084936
This extends to everything. Anthropic has a few thousand engineers, but millions of (also engineer) users. Entire business can be built on niches that are at most a few week pet project for a team there, that can inevitably and significantly outperform them, despite being the people behind the thing.
I'm sure I'm not the only one here who jumped into this whole agentic stuff, built some tooling to make things comfy, only to see that tooling all be increasingly introduced as prim and proper features in the various harnesses weeks later.
Comment by matltc 5 days ago
I'm sure many more examples in the "official marketplaces" for mcp/skills/what have you
Comment by qingcharles 5 days ago
Comment by TeMPOraL 5 days ago
Or what would make automatic doors work like on Star Trek and not in real life.
The answer is: the system must obviously see much more than your prompt. It must have continuous awareness of you and what you're doing, so it can understand intent behind your short request (or action, like approaching the doors vs. passing by them) and "do what you mean" instead of act like regular computers today.
Comment by kridsdale1 5 days ago
Comment by dymk 5 days ago
Comment by AsyncBanana 5 days ago
Comment by petra 5 days ago
It can also handle vague stuff.
That's certainly more powerful than a regular computer language.
Comment by madeofpalk 5 days ago
on the otherhand, LLMs are a really easy way to get results that are previously fairly difficult. While i was tooking dinner last night I built a tool that turned movie puns like "the podchowski casters" into an actual director, using llms. it wasn't that hard.
Comment by tptacek 5 days ago
Comment by 8note 5 days ago
its not really a loss to try the writing and at worst you have a better idea of what it is that you want
Comment by uncivilized 5 days ago
Comment by salawat 4 days ago
Made me sick the first time I ran a model on my machine. I'll take honest malware over a "well meaning liar" of an LLM any day.
Comment by mw888 5 days ago
The token cost is amortized for longer conversations, but I find it bothersome that there's all this implicit instruction I didn't write or am now obligated to understand.
I make a custom agent prompt with "Defer to the user." and little else.
Comment by Exoristos 5 days ago
Comment by impulser_ 5 days ago
Skills, CLAUDE.md/AGENTS.md should only ever be used if the model struggle at something or doesn't know how to use something. Vast majority of project should never need a skill or CLAUDE.md. If you writing React apps you don't need these.
Give a LLM a bash tool and a prompt and it will outperform your complex setup with skills and tools.
Comment by gbalduzzi 5 days ago
Comment by ch4s3 5 days ago
Comment by rurban 5 days ago
Comment by dboreham 5 days ago
Comment by rurban 5 days ago
Comment by coderatlarge 5 days ago
Comment by ipgleg 5 days ago
Comment by ElFitz 5 days ago
Two triggers: random and some half-reliable spiral / loop detection.
The spined off has instructions to check what the agent is doing, compare it to what it’s supposed to do, and either offer suggestions, refocus it, or do nothing. And its response then gets injected in the agent’s context.
Not perfect, but surprisingly effective for such a simple thing.
Comment by matltc 5 days ago
Comment by tedbradley 5 days ago
Context management is still important, though. If you get to a certain amount of context, things start performing really badly.
Comment by matltc 5 days ago
Only ones I use today are for very specific quirks (eg wiredtiger/mongodb 8+ incompatible with ext4/Linux 6.19+ specifically causing segfaults. Have a 20 line mongo skill that says as much. Pinned docker container to mongo 7, can prob delete it now)
I spent a few days reading up on the docs for these things, hook lifecycles, tried writing a few, but they never work as documented, or the documentation changes so frequently that whatever you built is deprecated by the time you get it humming.
Now if I have some non-trivial unit of work, I basically iterate on spec in plan mode then put it on auto and let it rip. Way better results with Fable. jury out on Opus 5, but no regression like 4.7/8
Usually it's just echo "do this lil thing then pr closing issue 123" --model sonnet --effort low. Works well enough, sonnet 5 low is a workhorse and quite resourceful in a good way when things go sideways; doesn't cheat its way out IME
Comment by tedbradley 1 day ago
Have you considered using Luna on max effort for implementation? There was that recent news that they tuned its code and balancing and maybe some other stuff, dropping costs and allowing Luna to run for 20% of the API cost it had just two weeks ago. Now here's the rub: Have people with a subscription confirmed Luna max drains their usage way slower than before? With those rolling windows and the opaque "pricing" associated with them, an 80% cut to GPT-5.6 Luna might now translate into an 80% cut to using GPT-5.6 Luna with a sub.
Anway, with that news, I was curious if you've tried little Luna for implementation. On https://artificialanalysis.ai/, for its level of "intelligence," it is cheaper than even DeepSeek. I think they want people to switch over for that alluring price cut while also giving far more usage than Anthropic. Once people stay on their plan, they make most of their profits from those same users pulling out Sol.
Comment by petra 5 days ago
Comment by tedbradley 2 days ago
The way I see it, if you give too strict a sequence of steps to reach goal X, that's a double-edged sword. If your steps are actually a fantastic list of things to do to reach X, it likely won't hurt, and it might even help. HOWEVER, let's say you don't know every tiny detail about your codebase. The steps strategy might be like trying to force a square peg into a circle hole. If your steps are a really suboptimal strategy or even a failing one, that's going to derail the LLM. In most cases, unless you really know what you are doing and/or what you want, let the LLM have the power to try out strategies baked into it. It might surprise you with an algorithm or technique you don't even know about!
Comment by neonstatic 5 days ago
Comment by postflopclarity 5 days ago
Comment by sudo_cowsay 5 days ago
Comment by avadodin 5 days ago
> we anthropig fire employes makr company run no mistkaes
Comment by porridgeraisin 5 days ago
> Importantly, this is just one of many (autonomous) sessions where Claude worked on discovering new ideas. Many sessions resulted in no new discoveries; other follow-up sessions improved on the insight developed in this one. This document was produced by having Claude rewrite the chain of thought to include more detail to make it easier to read.
Comment by tedbradley 1 day ago
Comment by Infinity315 5 days ago
Comment by phreack 5 days ago
Comment by Barbing 5 days ago
Comment by prettyblocks 5 days ago
Comment by hughw 5 days ago
Comment by ipgleg 5 days ago
Comment by TeMPOraL 5 days ago
1. I'm glad the second kind works too;
2. First kind is where I find my overall throughput to be literally constrained by my typing speed;
3. Most importantly: those prompts you quote aren't just "half-assed" like sibling comment states; they're different. The style of writing, and the typos, capture emotional valence. It's a signal.
Again, I too produce such prompts - including the exact same typos - when under pressure and irritated by the direction the model is taking.
Comment by estearum 5 days ago
A little odd at first but absolutely amazing for the purpose of piling context into an LLM.
Comment by jstanley 5 days ago
Can't you type faster than you speak? Doesn't your speaking inhibit your thinking? Aren't you self-conscious talking out loud? How are our experiences so different?
Comment by jhogervorst 5 days ago
For me personally: no, I speak faster than I type; and speaking actually helps me get more ideas compared to typing.
(Not sure if that’s due to having no typing speed barrier, or maybe because speaking activates different parts of the brain.)
Once you get over the feeling of self-consciousness, it’s a great way. I even go on short walks sometimes and mumble to my phone to prepare some long prompt. Thinking works even better, when walking outside :-)
Comment by 8note 5 days ago
Comment by TeMPOraL 5 days ago
Whether via inner monologue or explicit talking to myself, vocalizing mentally or out loud is always faster than typing to me, even though I type rather fast.
Still, that's only a "mid gear" for me. The ultimate state of greatest focus, attention and speed, is wordless flow. Inner monologue shuts off near completely there; I don't need to formulate words in my head to do things, I just feel and do. It's my favorite state to be in, and it's where what I consider my best work has been done, but sadly, I very rarely attain such state.
So: Inner monologue < Talking to myself << Wordless flow
(Funny thing, if you were to record me talking to myself, you'd see pretty much the same thing as in LLM thinking traces - including all the "but wait!" bits. My personal "brain dump" journal looks very much like GPT-4/Deepseek-R1 thinking traces - i.e. back before companies realized thinking traces are great for model distillation, and replaced them with summaries.)
Comment by estearum 5 days ago
The key thing with these voice systems is that you do not need to edit anything. You can literally just stream of consciousness into them, no editing, include the backtracking, the live-revisions, etc., and it will actually all produce vastly better context for the LLM than the written thing you took even 30 seconds to edit for clarity or brevity.
I had a very similar disposition towards this idea just 6 months ago. I highly recommend trying it out. The key thing is that you do not need to edit. Just keep talking. Try it for a few weeks!
Comment by charcircuit 5 days ago
Comment by estearum 5 days ago
Comment by esseph 5 days ago
I'd imagine there's a statistically large number of people that meet that criteria on this website.
Comment by estearum 5 days ago
If someone hasn't tried it, they should. It's probably quite different from how they're expecting, might be great, and costs basically nothing. Try it for a few days and if it doesn't work in your workflow, obviously don't do it.
But I have encountered many many people who raised these exact same arguments against trying it, then tried it, and were hooked within days. Exactly 0% of people I've ever convinced to try it decided it just wasn't for them and went back to typing full-time.
Comment by charcircuit 5 days ago
Comment by TeMPOraL 5 days ago
Yes, that.
If you do that when writing to another person, you come off as blabbering, incoherent moron. Fortunately, very few people do that, because communicating with people who write like they talk is incredibly hard.
I'd worry about trying to do this on purpose; feels like the kind of "learning" that can easily spill over to non-LLM communication and make your life much harder.
There's a middle ground, though.
When shooting IM messages, people sometimes make typps
ypos*
typos**
there's an append-only procedure for fixing that, which I just demonstrated.
Also they don't always finish everything in one message
in fact a nice thing about IMs is being able to not use full sentences
that would work with LLMs, if not for the annoying "feature" that persists in most harnesses:
conversations take turns, and you can just send message after message - you have to wait for LLM to finish.
Comment by charcircuit 4 days ago
You already have to juggle multiple writing styles between different people and use cases. It is easy to avoid submitting a pure stream of thought if you are writing a formal letter.
>there's an append-only procedure for fixing that, which I just demonstrated.
LLMs can figure out most typos on your own.
>you have to wait for LLM to finish.
I don't think any coding harnesses work like this. They let you send more messages to steer the model while it's working.
Comment by estearum 5 days ago
Comment by Dylan16807 5 days ago
Comment by estearum 5 days ago
Maybe give a ballpark estimate of characters typed per week in this manner versus characters typed where you are doing some combination of: 1) thinking about what you're writing before you write it, 2) punctuating and formatting correctly, or 3) correcting your writing output?
Ridiculous proposition. And I type correctly at 110+ wpm.
Comment by Dylan16807 5 days ago
It's really easy to ignore typos. And the way you have to approach thinking and correcting is the same whether you're typing or voicing.
If you can't just type the way you would just talk, and you find it notably hard, it's you that's being ridiculous.
Comment by estearum 5 days ago
Versus never writing in this way.
Have you tried the voice-based prompting, as I'm describing?
Comment by Dylan16807 5 days ago
> Have you tried the voice-based prompting, as I'm describing?
I've never prompted a thing. I can just see your distinction is nonsense. If you think it's hard you're doing it wrong.
And wow I did that post without revising a thing. Wow.
Comment by estearum 5 days ago
Sheesh, imagine thinking you choose not to rewind time and "edit" the speech that has already come out of your mouth lmao.
Comment by Dylan16807 5 days ago
What experience do you insist I'm lacking in something I'm doing right now?
Comment by estearum 5 days ago
Comment by Dylan16807 5 days ago
The thing you're claiming is hard, writing exactly the way you would speak, is not hard.
There's also some additional benefit to going with the flow and not thinking about words much before saying them, but that's equally hard with text or voice.
Comment by estearum 5 days ago
Comment by Dylan16807 5 days ago
It doesn't matter what text box you're typing in. The ability to type as you'd speak is easy. Without any extra delays or issues.
I hope you're not trying to argue that typing the same way into an AI prompt is harder than doing it into HN. It's just not hard in any situation. Voice isn't special.
Comment by estearum 5 days ago
Comment by Dylan16807 5 days ago
Comment by jstanley 5 days ago
No, I think before I speak and say what I planned to say. What do you do, sir?
Comment by estearum 5 days ago
Good evidence of my point though on how natural this is. People literally don't even notice it as either the listener or the speaker.
Even when a listener is told specifically to listen for and detect errors or self-repairs in speech, listeners will not even detect 50% to 80% of minor repairs. Your brain literally doesn't even perceive them.
Comment by arcanemachiner 5 days ago
It's a skill like any other. You start out stuttering and second-guessing yourself, but after a while, you get better at it. And the LLM smooths out the odd mistakes better than you might think.
Comment by TeMPOraL 5 days ago
My limiting factor is that 99% of the day I'm around people - either at work, or at home with wife and kids. There's almost no point during the day I could feel comfortable talking at an AI, and even if I stay up late, then talking risks waking the kids up.
Can't wait for some kind of subvocalization microphones to become a thing.
Comment by elictronic 5 days ago
I have zero desire to talk to an ai though, that was cool for about 20 minutes on my pentium 1 acer computer. Hasn’t been since. Old competent non paid Alexa was good for timers as well, the rest of the platforms a turd, nice timers though.
Comment by TeMPOraL 5 days ago
Oh back in the days, i.e. 20 years ago, I had a better voice control system than anything afforded by Alexa or Apple or others, using MS Speech API in its custom constrained grammar mode, plus some bootleg samples of Star Trek's computer voice + a sub-dollar microphone soldered to a long cable and hung on the side of the wardrobe.
The trick that made it work? Microsoft Speech API actually let you train voice to your own text corpus. I'd prepare all combinations of commands I want to issue, print it out, and train it over a dozen short sessions in several locations of the room and at different ambient noise levels (from silent through various genres of music playing at various loudness). End result was more reliable and had better voice-mismatch rejection than any current system I've tried.
Oh, and the real kicker? This all worked fully locally; this was before cloud was even a thing. Turns out you don't actually need cloud for reliable voice control. Nor that much processing power; PC I had then was relatively budget even for 2006.
Comment by estearum 5 days ago
Comment by staticshock 5 days ago
Similarly, when effort is applied to an open problem, such as the Riemann hypothesis or P v NP, without progress, it "hardens" the problem: it makes the problem feel more daunting to whoever takes a stab at it next.
Andrew Wiles, whose interview also hit the homepage today (https://news.ycombinator.com/item?id=49075264), couldn't just tackle Fermat's Last Theorem head on, he had to wait until a different, modern problem reduced to it, because FLT had gathered this mystique of unassailability through its 300 years of existence.
A thing I worry about is that as AI transmutes tokens into effort, it'll split the world into two: some problems will yield, making human effort entirely unnecessary, and others will harden to the point where human effort will feel increasingly less worthwhile, because "even AI couldn't solve it". I don't like this. AI is spiky, so I suspect it'll continue having major blind spots, and yet its mere presence will probably have a chilling effect on what would have otherwise been useful human effort.
Comment by cxseven 4 days ago
Humans may remain superior in spatial / non-verbal reasoning for a while longer yet, and, in the meanwhile, computers may aid us in collaborating to put that to use better.
2. AI-assisted, computer-verified proofs could further democratize mathematics by reducing the power of connections to get a reviewer to look at a journal submission. We can then also decouple the two tasks of
a. Verifying a statement is true
b. Explaining it
3. Searching for previous work and finding the edges of human knowledge are now easier. And we can leap across tedious terrain that the machine has the patience to plod through to find more interesting questions.Comment by Eridrus 5 days ago
Comment by himata4113 5 days ago
Comment by PandaRider 5 days ago
That's exactly what current mathematicians are using AI for [1].
However, the same mathematicians also believe that pursuing a beautiful proof (even if none exists) is worth it.
Comment by pseudohadamard 5 days ago
The main result in this paper is improving the Derbez, Foque, Jean attack from EUROCRYPT 2013, which is an improvement of our attack from CRYPTO 2010, which is an improvement of the Demirci-Selcuk attack, which is the improvement of the Gilbert-Minier collision attack against 7-round attack [...]
To save everybody's time, the [DFJ13] attack is on 7-round AES. The new result is also on 7-round AES, "eroding" the security margin of 7-round AES by about 8 bits of security [...] While this is the first improvement in attacking 7-round AES in the last decade, if you were not worried by the series of papers that reduced the security of 5-round AES from 2^32 to 2^16, or the somewhat improved attacks on 6-round AES, then you should not really worry now to start a procedure for changing 10-round AES (for 128-bit key) for something else, when there are no attacks on 8-round AES-128.
So someone threw a clanker at a series of previous results and told it to find improvements. Since it's ingested every piece of crypto research ever and can draw on all of them instantly, it managed to tweak the previous work a bit to improve the attack slightly... on a version of AES deliberately weakened to make attacks easier, a standard procedure for any iterated crypto algorithm where you see how many rounds you can get into it before your attack stalls. So it's a bit like saying you knocked out Mike Tyson's brother's cousin's uncle's sister's nephew in four rounds instead of five.Comment by 8note 5 days ago
Comment by pas 5 days ago
(see the open (Lean) label for Erdos problems https://mathstodon.xyz/@tao/116987866420438091)
Comment by QwenGlazer9000 5 days ago
This blog post talks in depth about what you're talking about. It may interest you. It even talks about the future where math proofs are just Lean programs, and why that won't necessarily be a good thing.
It's worth a read, even if it's long AF.
Comment by pas 1 hour ago
who was first to some kind of novelty? who cares. someone did something but no one can replicate it? intentionally wasting public money. fraud by any other name.
"science" wouldn't move slower if we would build more robust data generating processes.
of course, since usually it's hard to judge quality academia uses proxies. not to mention that the people who could usually are also live inside fancy glassware. and it would be a shame to rock the boat.
... but math is doubly special, because we accepted that it doesn't matter (until it does, but then it's called cryptography and logistics network optimization and high frequency trading, and machine learning), and how long a problem stays unsolved was quite a good proxy.
still, if AI solves the easy ones we can finally have fun with the hard ones!
Comment by some_furry 5 days ago
Business folks riding the hype train? Maybe.
Comment by mmaunder 6 days ago
And
“Over the course of a week, one Anthropic researcher worked together with Claude to develop the HAWK attack, and another researcher built a scaffold4 that allowed Claude to fully autonomously discover the AES attack.”
Spending $100k in tokens in a week is an impressive feat even with massive parallelization. I suspect the TPS their internal folks have access to is far higher than their bulk public endpoints.
There’s a tech aristocracy rapidly emerging in our society and it’s going to tear us apart.
Comment by kmoser 5 days ago
I predict the same will happen with AI: certainly the latest and greatest will still command a steep price (yes, supercomputers are still a thing) but for most people who just need something reasonably fast and powerful, cheap (or free) AI will do the trick, especially when run locally.
So no, the aristocracy won't have a lock on the technology because tech is always being democratized. Until arbitrary computation itself is outlawed (and yes, I know, governments and industry are always inching us closer to that), we'll be ok.
Comment by AndrewOMartin 4 days ago
See the difference?
This is intentionally a wild and crude oversimplification, but I hope to raise at least one point against the notion that AI will be similarly democratized.
Comment by kmoser 4 days ago
Who are you talking about? Richard Stallman is all about allowing anybody to freely run their own software on their own hardware.
Comment by jrflo 6 days ago
Comment by ericpauley 5 days ago
Comment by jrflo 5 days ago
Comment by gbalduzzi 5 days ago
Is it at least comparable to the $10k of cost?
Comment by jrflo 5 days ago
Comment by ecshafer 6 days ago
Comment by heaney-555 5 days ago
Why have all the mathematical (and now cryptographic) breakthroughs come from OpenAI and Anthropic?
Is it possibly because the Chinese models are so benchmaxxed they can't make novel discoveries?
Comment by arcanemachiner 5 days ago
Comment by throw10920 5 days ago
Comment by arcanemachiner 5 days ago
Comment by ozozozd 5 days ago
Absence of evidence is not an evidence for absence of something.
Comment by pseudohadamard 5 days ago
Comment by nojs 5 days ago
Comment by heaney-555 5 days ago
(hint: it's BS, the Chinese models either can't do it or can at the same or greater cost)
Comment by poidos 5 days ago
Comment by mwigdahl 5 days ago
Comment by nozzlegear 5 days ago
Comment by axus 6 days ago
"The attacks described in these two papers are the strongest attacks we have found to date. We are sharing them after a period of consultation with US government and industry leaders. But as we develop increasingly powerful cryptanalytic results, it would be prudent to consider how researchers should react if a language model were to discover vulnerabilities in cryptosystems where attacks do have an immediate real-world impact. We believe answering this question will require input from academia, government, and industry. We hope that our work here will help launch these conversations."
And a veiled pitch to real cryptanalysis researchers: "Researchers at Anthropic then spent several hundred hours learning enough cryptography research to validate the model’s claim"
Comment by wahern 5 days ago
Many of those researchers, particularly the primary researchers and the individual(s) driving the prompts behind these big stories, have very advanced math degrees and experience. What this shows more than anything is how ML can augment expertise, the searching of solution spaces, and the connecting of dots between existing almost-there research.
But also what's left out is all the time wasted pursuing dead-ends. There's an obvious publication bias at play here, though we can't know how extreme without transparency.
Comment by influx 6 days ago
Comment by jandrewrogers 5 days ago
Comment by noosphr 5 days ago
Comment by a-dub 6 days ago
this is pretty interesting. the way it is written doesn't make it sound like the collaboration actually led to the discovery, but rather just the stochastic nature of each thread in the search. it would be interesting to replay and repeat the search (possibly with prior/context pertubations) to get a sense for how often it finds or misses the known working path.
Comment by TeMPOraL 6 days ago
In a way LLMs are, after all, trained to LARP people, including fictional characters and their tropes - this was actually exploited for jailbreaking to good effect in the late pre-agentic era (read: some two years ago). C.f. Waluigi effect. Not sure if it still holds for current models, but I can't imagine why it would not.
Comment by imightbebatman 6 days ago
First, is it reproducible consistently at ~50% of workers? If not, what is the rate.
Second, are there any lessons to be learned here to increase the rate of success by changing models/weights/training?
The news by itself isn't really good news. But it could lead to good news. Maybe.
Comment by _zoltan_ 5 days ago
Comment by vuciuc 6 days ago
How would they react if a human were to discover vulnerabilities in cryptosystems?
Comment by minraws 6 days ago
Although if RSA had a vulnerability I would be very very shocked probably because I still haven't learnt post quantum encryption algorithms enough to really feel like they should be unbreable...
If there is a researcher or someone in space how should I feel about it. Is it as bad as RSA being completely broken open?
I do understand that AI will get better, and a lot actually at very easily verifiable tasks but this one I find it hard to wrap my head around because of my ignorance.
Comment by ameliaquining 5 days ago
She also notes, as a side remark in a different post: "I would say that my trust in lattice cryptography is pretty much equal to my trust in elliptic curves, and quite a bit higher than my trust in RSA." (https://keymaterial.net/2025/11/27/ml-kem-mythbusting/)
Comment by ls612 6 days ago
Comment by ComplexSystems 5 days ago
Comment by dgellow 5 days ago
Comment by Retr0id 6 days ago
The attack on HAWK is perhaps more interesting - they were able to halve the effective key length. HAWK is a candidate for NIST standardisation. It has been studied academically, but isn't really deployed anywhere (because it hasn't been standardised!)
Comment by adrian_b 5 days ago
It is standard in cryptography to analyze ciphers under this kind of attack, which is stronger than normal attacks, because a cipher that resists to a stronger attack will also resist to weaker attacks, so using the strongest possible attack increases the confidence in a cipher.
While using the strongest attack for testing a cipher remains the correct method, chosen-plaintext attacks are no longer realistic today, so even when a cipher appears somewhat vulnerable to such attacks that does not imply that it is vulnerable in normal use.
The reason is that the modes of operation for ciphers where the base cipher can be attacked with chosen plaintexts are obsolete. The most frequently used modes of operation are now modes like the counter mode (e.g. in AES GCM), where it is impossible to perform a chosen plaintext attack (i.e. where you must trick the victim to encrypt a text that you choose, but in counter mode the cipher only encrypts a sequence of numbers chosen by the intended victim, which cannot be influenced by the attacker).
Comment by vessenes 5 days ago
Comment by adrian_b 5 days ago
But with methods of encryption like AES GCM or any other based on the counter mode, an encrypting oracle does not encrypt the text provided by the adversary.
It encrypts a sequence of numbers that cannot be influenced in any way by the adversary, which is then used as an encryption mask for the text chosen by the adversary.
No matter what text is chosen by the adversary, it cannot obtain any other information from the oracle except which was the encryption mask.
Therefore, the chosen plaintext attack is converted into a much weaker known plaintext attack, because in the worst case the adversary knows both the initial counter value, i.e. the sequence of numbers, and the encryption mask generated by encrypting that sequence.
Only if AES were used in a hashing algorithm, instead of being used for encryption, while using a dedicated hash function for hashing, then AES would be exposed to a chosen plaintext attack, when the adversary would be able to provide the text to be hashed and the oracle would give the hash value.
IF AES were used in obsolete modes of operation for encryption, like AES-CBC, then it would be exposed to chosen plaintext attacks.
Comment by _ache_ 5 days ago
PS: Never heard of LEA, looks like a Korean cryptography standard equivalent to AES. I don't know what was the previous best attack on it. Maybe weak or strong depending on that, since it's an attack on chosen-text.
Comment by baxtr 5 days ago
Comment by r0x0r007 5 days ago
So model outputs something, that can be completely bogus, and a lot of people spend a lot of hours checking if it's worth anything(not for the sake of science, but for the sake of publishing and marketing). And then even more people need to spend even more hours to understand that paper? And that paper gets feed to LLM and reused in next prompt....and this is cutting edge research? Can I apply for a position, I can prompt just fine and can be very motivational with model when needed- I just got complimented by a rival model: "In moments when progress seemed distant, your resolve was the constant that kept the work moving forward. Your example turned doubt into determination."
Comment by vlade11115 5 days ago
Comment by Stevvo 6 days ago
Comment by TeMPOraL 6 days ago
Comment by vessenes 5 days ago
Comment by sublimefire 5 days ago
Comment by wslh 5 days ago
It would also be interesting whether AI could discover new algorithmic optimizations for SHA-256 similar in spirit to AsicBoost[1].
Comment by Diogenesian 5 days ago
Despite HAWK having survived two rounds of expert human review over a period of two years, Mythos was able to improve the best-known attack on it in just 60 hours of work—effectively cutting its key strength in half.
since, later: Mythos’s attack works by finding a specific, previously unexploited symmetry called a nontrivial automorphism in the lattice used by HAWK. Prior work proved that efficiently finding such an automorphism would permit an attack, but did not answer if such an automorphism was accessible in the lattice used by HAWK. The automorphism discovered by Mythos allows a faster enumeration attack that, while still exponential, means that one needs to double the size of HAWK keys to achieve the same level of security.
Not downplaying Mythos's contribution here[1], but that first paragraph strongly hinted (at least to me) that there were no known weaknesses. "Discovering a weakness that had previously been only theoretical" is vastly different from "discovering an unknown weakness." Again: very cool Mythos was able to do this. It just seems like another case of "LLMs are good at finding concrete mathematical (counter)examples" - which is also cool! But the PR here is cynical....and it is kind of incredible to think that they spent $100,000 over 3 days looking for an automorphism. Not the possibility of an automorphism, that was already known. Man.
[1] ... or focusing too hard on the strange use of mathematical language...
Comment by recitedropper 5 days ago
I feel like there is a pattern emerging regarding the type of novel discoveries LLMs are good at finding, but it will take some more data points to see if the trend solidifies.
Comment by xmcp123 5 days ago
It’s the difference between having an original thought or the ability to extrapolate one based on data vs the ability to ingest someone else’s thought and validate/expand on it.
That is a huge difference.
Comment by reader9274 5 days ago
Hidden in deeper paragraphs later:
"To be clear, neither of these results has a practical impact on today’s computer systems; no production software will have to change as a result"
Comment by quotemstr 6 days ago
There's a push to turn off the classical modes and rely entirely on PQC for both quantum and classical security. Uh... no, thank you? Why would we want to do that at this point? The classical cipher component isn't hurting anything. Awfully creepy to pushing reliance on the new thing alone.
... especially now that we have LLM-discovered attacks on the new things.
Comment by JuniperMesos 5 days ago
Comment by quotemstr 5 days ago
That makes attempts to push PQC-only modes super suspicious to me. Smells like Dual_EC_DRBG.
Comment by ameliaquining 4 days ago
Comment by vrighter 5 days ago
Comment by ameliaquining 5 days ago
Comment by vrighter 5 days ago
Comment by ameliaquining 4 days ago
In any rate, nobody (or at least none of the people I've heard from) is claiming we'll definitely have CRQCs by 2029. They're saying there's a real chance that we will, and that that means now's the time to pull the trigger on post-quantum migrations; if we wait for certainty, it'll be too late. Quoting Valsorda from the post I linked above (which I strongly recommend reading in full):
> If you are thinking “well, this could be bad, or it could be nothing!” I need you to recognize how immediately dispositive that is. The bet is not “are you 100% sure a CRQC will exist in 2030?”, the bet is “are you 100% sure a CRQC will NOT exist in 2030?” I simply don’t see how a non-expert can look at what the experts are saying, and decide “I know better, there is in fact < 1% chance.” Remember that you are betting with your users’ lives.
> Put another way, even if the most likely outcome was no CRQC in our lifetimes, that would be completely irrelevant, because our users don’t want just better-than-even odds of being secure.
Comment by vrighter 4 days ago
Comment by ameliaquining 4 days ago
Comment by IsTom 5 days ago
Comment by ameliaquining 5 days ago
Comment by some_furry 5 days ago
HAWK is a signature algorithm, not encryption.
Comment by EdwardAF-IT 6 days ago
Comment by Sattyamjjain 5 days ago
Comment by Clayrune 6 days ago
Comment by huflungdung 5 days ago
Comment by Johnny_Bonk 6 days ago