Discovering Cryptographic Weaknesses with Claude

(anthropic.com)

116 points | by gslin 4 hours ago

12 comments

  • _dwt 2 hours ago

    I find that some of my friends and acquaintances have gotten obsessed with prompting style, "prompt engineering", which skills to use, which skills to build, "context engineering", and a billion other variations on "how to write smart things so the model does good".

    Friends, look at the prompts that Anthropic's own people are putting into the machine:

    > A few hours after the first message, we found that Claude was still searching for simple attacks and sent a message: “no again the goal is that we have highly inteligent [sic] model as good top researcher, we want to find new attacks”;

    > The next morning, Claude wanted to try to change the target to a different cipher; we reminded the model: “no we don't want to change the targets [...] agian [sic] we need to find something that worth [sic] publishing”;

    > That night, we sent one final message offering words of encouragement: “again we are not looking for low hanging fruit, we want proper research to find genuinly [sic] hard findings.”

    All of that RLHF and fine-tuning effort is going toward making prompts like this, or worse, work with no fuss.

    • impulser_ 1 hour ago

      Yeah, in fact you can actually make the model perform worst. You should allow the model to "think for itself" instead of pushing your reasoning into the prompt. You should give it simple prompt and steer it along the way.

      Skills, CLAUDE.md/AGENTS.md should only ever be used if the model struggle at something or doesn't know how to use something. Vast majority of project should never need a skill or CLAUDE.md. If you writing React apps you don't need these.

      Give a LLM a bash tool and a prompt and it will outperform your complex setup with skills and tools.

      • gbalduzzi 14 minutes ago

        You need agents.md and similar for indications about stuff that is not in the code itself. There are plenty of use cases for that, the alternative is have the LLM guess the most probable solution, which may be correct but may also be wrong.

        • rurban 52 minutes ago

          No, we still need CLAUDE.md to override bad system prompt goals. The system prompt goes for the simple fast solution, no error handling, no hardening, no abstractions, but overly verbose comments. Vast amount of people need to override this AI slop.

          • dboreham 42 minutes ago

            Never once had to do that, fwiw.

        • qingcharles 1 hour ago

          It's counter to what sci-fi taught us using AI would be like. We never thought we'd have to feed it words of encouragement, we expected it to act more mechanically, like the computer interfaces we have been using, but here we are. It's kind of quaint, and kind of endearing.

          • kridsdale1 11 minutes ago

            C-3PO was neurotic and needed lots of reassurance.

            • TeMPOraL 1 hour ago

              Been watching the wrong sci-fi :). Star Trek had both AI modes as primary characters, in form of the ship's computer, and Data (TNG) / the Doctor (VOY). LLMs are actually great at acting as both, but I don't believe people thought much about what is required to make the "simpler" interaction mode of the ship's computer to work.

              Or what would make automatic doors work like on Star Trek and not in real life.

              The answer is: the system must obviously see much more than your prompt. It must have continuous awareness of you and what you're doing, so it can understand intent behind your short request (or action, like approaching the doors vs. passing by them) and "do what you mean" instead of act like regular computers today.

            • perching_aix 17 minutes ago

              To quote - what I found to be - an absolute zinger from another trending thread on here just 6 hours ago:

              > Typical users run software written by atypical users.

              https://news.ycombinator.com/item?id=49084936

              This extends to everything. Anthropic has a few thousand engineers, but millions of (also engineer) users. Entire business can be built on niches that are at most a few week pet project for a team there, that can inevitably and significantly outperform them, despite being the people behind the thing.

              I'm sure I'm not the only one here who jumped into this whole agentic stuff, built some tooling to make things comfy, only to see that tooling all be increasingly introduced as prim and proper features in the various harnesses weeks later.

              • TeMPOraL 2 hours ago

                Having spent a lot of time over the last months writing both refined, structured, and grammatically perfect prompts, as well as exact equivalents of the ones you quoted (modulo subject of the prompt), I have three observations:

                1. I'm glad the second kind works too;

                2. First kind is where I find my overall throughput to be literally constrained by my typing speed;

                3. Most importantly: those prompts you quote aren't just "half-assed" like sibling comment states; they're different. The style of writing, and the typos, capture emotional valence. It's a signal.

                Again, I too produce such prompts - including the exact same typos - when under pressure and irritated by the direction the model is taking.

                • estearum 2 hours ago

                  Have you tried using Wispr or Willow (or any one of a thousand alternatives?)

                  A little odd at first but absolutely amazing for the purpose of piling context into an LLM.

                  • jstanley 1 hour ago

                    It's so bizarre to me that people want to do this.

                    Can't you type faster than you speak? Doesn't your speaking inhibit your thinking? Aren't you self-conscious talking out loud? How are our experiences so different?

                    • jhogervorst 26 minutes ago

                      > Can't you type faster than you speak? Doesn't your speaking inhibit your thinking?

                      For me personally: no, I speak faster than I type; and speaking actually helps me get more ideas compared to typing.

                      (Not sure if that’s due to having no typing speed barrier, or maybe because speaking activates different parts of the brain.)

                      Once you get over the feeling of self-consciousness, it’s a great way. I even go on short walks sometimes and mumble to my phone to prepare some long prompt. Thinking works even better, when walking outside :-)

                      • estearum 1 hour ago

                        These are all interrelated points and the sibling comment is correct: it's a skill.

                        The key thing with these voice systems is that you do not need to edit anything. You can literally just stream of consciousness into them, no editing, include the backtracking, the live-revisions, etc., and it will actually all produce vastly better context for the LLM than the written thing you took even 30 seconds to edit for clarity or brevity.

                        I had a very similar disposition towards this idea just 6 months ago. I highly recommend trying it out. The key thing is that you do not need to edit. Just keep talking. Try it for a few weeks!

                        • charcircuit 16 minutes ago

                          You don't need to edit anything when you are typing either. You don't even need to worry about spelling or typos.

                        • arcanemachiner 1 hour ago

                          My stream of barely-intelligible rambling comes out much faster than I can type, and it's not even close.

                          It's a skill like any other. You start out stuttering and second-guessing yourself, but after a while, you get better at it. And the LLM smooths out the odd mistakes better than you might think.

                        • TeMPOraL 2 hours ago

                          Voice dictation tools? Not really. Tried various dictation tools interacting on the phone, but the quality varies, and resulting prompts are very much not like I would write them.

                          My limiting factor is that 99% of the day I'm around people - either at work, or at home with wife and kids. There's almost no point during the day I could feel comfortable talking at an AI, and even if I stay up late, then talking risks waking the kids up.

                          Can't wait for some kind of subvocalization microphones to become a thing.

                          • estearum 1 hour ago

                            I highly recommend getting a boom mic that you can keep right up against your mouth and literally whisper into. Take a few weeks just doing stream-of-consciousness prompting into it like this. Don't worry about editing or backtracking or revisions. It's really amazing how well the AI voice recognition → coding agent workflow works.

                            • elictronic 2 hours ago

                              Get a noise machine for wife and kids bedrooms. I talk to friends late at night and my deep voice carries through walls. Works great.

                              I have zero desire to talk to an ai though, that was cool for about 20 minutes on my pentium 1 acer computer. Hasn’t been since. Old competent non paid Alexa was good for timers as well, the rest of the platforms a turd, nice timers though.

                              • TeMPOraL 1 hour ago

                                > that was cool for about 20 minutes on my pentium 1 acer computer. Hasn’t been since.

                                Oh back in the days, i.e. 20 years ago, I had a better voice control system than anything afforded by Alexa or Apple or others, using MS Speech API in its custom constrained grammar mode, plus some bootleg samples of Star Trek's computer voice + a sub-dollar microphone soldered to a long cable and hung on the side of the wardrobe.

                                The trick that made it work? Microsoft Speech API actually let you train voice to your own text corpus. I'd prepare all combinations of commands I want to issue, print it out, and train it over a dozen short sessions in several locations of the room and at different ambient noise levels (from silent through various genres of music playing at various loudness). End result was more reliable and had better voice-mismatch rejection than any current system I've tried.

                                Oh, and the real kicker? This all worked fully locally; this was before cloud was even a thing. Turns out you don't actually need cloud for reliable voice control. Nor that much processing power; PC I had then was relatively budget even for 2006.

                        • neonstatic 2 hours ago

                          > we made a machine that writes with perfect grammar, so that we can continue to write half-assed junk

                          • postflopclarity 2 hours ago

                            claude's grammar has gotten significantly worse with latest versions. it no longer writes perfectly at all

                            • sudo_cowsay 2 hours ago

                              It's called being more human-like lol XD

                          • porridgeraisin 2 hours ago

                            To be fair

                            > Importantly, this is just one of many (autonomous) sessions where Claude worked on discovering new ideas. Many sessions resulted in no new discoveries; other follow-up sessions improved on the insight developed in this one. This document was produced by having Claude rewrite the chain of thought to include more detail to make it easier to read.

                          • staticshock 1 hour ago

                            When high quality effort is applied to a tool, such as AES or the linux kernel, we intuit that it "hardens" the tool. That is, it makes the tool more correct, more resilient, less assailable, etc.

                            Similarly, when effort is applied to an open problem, such as the Riemann hypothesis or P v NP, without progress, it "hardens" the problem: it makes the problem feel more daunting to whoever takes a stab at it next.

                            Andrew Wiles, whose interview also hit the homepage today (https://news.ycombinator.com/item?id=49075264), couldn't just tackle Fermat's Last Theorem head on, he had to wait until a different, modern problem reduced to it, because FLT had gathered this mystique of unassailability through its 300 years of existence.

                            A thing I worry about is that as AI transmutes tokens into effort, it'll split the world into two: some problems will yield, making human effort entirely unnecessary, and others will harden to the point where human effort will feel increasingly less worthwhile, because "even AI couldn't solve it". I don't like this. AI is spiky, so I suspect it'll continue having major blind spots, and yet its mere presence will probably have a chilling effect on what would have otherwise been useful human effort.

                            • Eridrus 27 minutes ago

                              This is a problem that will solve itself, people will continue to work on the problems that AI fails at, likely by telling AI the approaches they want AI to take.

                              • some_furry 35 minutes ago

                                I wouldn't worry about too many mathematicians adopting the "even AI couldn't solve it" attitude.

                                Business folks riding the hype train? Maybe.

                              • mmaunder 3 hours ago

                                “Each of the results cost roughly $100,000 in API cost to develop.”

                                And

                                “Over the course of a week, one Anthropic researcher worked together with Claude to develop the HAWK attack, and another researcher built a scaffold4 that allowed Claude to fully autonomously discover the AES attack.”

                                Spending $100k in tokens in a week is an impressive feat even with massive parallelization. I suspect the TPS their internal folks have access to is far higher than their bulk public endpoints.

                                There’s a tech aristocracy rapidly emerging in our society and it’s going to tear us apart.

                                • kmoser 1 hour ago

                                  Today's state-of-the-art AI systems are the equivalent of late 1970s personal computers: bulky and expensive, but wildly more powerful than what came before. And look what happened: tech improved by orders of magnitude, and eventually computers were tiny and cheap.

                                  I predict the same will happen with AI: certainly the latest and greatest will still command a steep price (yes, supercomputers are still a thing) but for most people who just need something reasonably fast and powerful, cheap (or free) AI will do the trick, especially when run locally.

                                  So no, the aristocracy won't have a lock on the technology because tech is always being democratized. Until arbitrary computation itself is outlawed (and yes, I know, governments and industry are always inching us closer to that), we'll be ok.

                                  • jrflo 3 hours ago

                                    That's not really that ridiculous. Looking at my ChatGPT stats my biggest day of token usage was 1B tokens (seeing how far Sol Ultra could go on a difficult problem with a quantitative goal and eval harness that it could run on it's own that allowed it to keep going until it succeeded). I blew through my $100 subscription usage in that one day, but with a 80/20 token blend that's $10k in API billing. So, $70k in a week. With the higher cost of Mythos, that's not that crazy. It's more so a testament to how ridiculously marked up API tokens are, or how discounted subscriptions are (who knows which is true)

                                    • gbalduzzi 8 minutes ago

                                      How much was valuable the output you got during the day?

                                      Is it at least comparable to the $10k of cost?

                                    • ecshafer 3 hours ago

                                      So $1-10k in Chinese model time, thus why we must ban them.

                                      • heaney-555 1 hour ago

                                        If a Chinese model can do it for $1-10K, then why hasn't one?

                                        Why have all the mathematical (and now cryptographic) breakthroughs come from OpenAI and Anthropic?

                                        Is it possibly because the Chinese models are so benchmaxxed they can't make novel discoveries?

                                        • arcanemachiner 1 hour ago

                                          They're a few months behind the frontier, which, by the way, only started making these discoveries in the last few months.

                                          • poidos 1 hour ago

                                            Don't know if one has or not, but a lack of announcement is not a lack of success.

                                            • mwigdahl 1 hour ago

                                              Do you really think that if a Chinese model had achieved a significant math breakthrough that it wouldn't be trumpeted to the global media? The "Deepseek moment" was great for China; this would be the same.

                                      • axus 3 hours ago

                                        I can already picture the faces of national security directors everywhere.

                                        "The attacks described in these two papers are the strongest attacks we have found to date. We are sharing them after a period of consultation with US government and industry leaders. But as we develop increasingly powerful cryptanalytic results, it would be prudent to consider how researchers should react if a language model were to discover vulnerabilities in cryptosystems where attacks do have an immediate real-world impact. We believe answering this question will require input from academia, government, and industry. We hope that our work here will help launch these conversations."

                                        And a veiled pitch to real cryptanalysis researchers: "Researchers at Anthropic then spent several hundred hours learning enough cryptography research to validate the model’s claim"

                                        • wahern 10 minutes ago

                                          > And a veiled pitch to real cryptanalysis researchers: "Researchers at Anthropic then spent several hundred hours learning enough cryptography research to validate the model’s claim"

                                          Many of those researchers, particularly the primary researchers and the individual(s) driving the prompts behind these big stories, have very advanced math degrees and experience. What this shows more than anything is how ML can augment expertise, the searching of solution spaces, and the connecting of dots between existing almost-there research.

                                          But also what's left out is all the time wasted pursuing dead-ends. There's an obvious, very extreme publication bias at play here.

                                          • influx 3 hours ago

                                            It would be shocking if they haven't been pulling on these threads for as long as they've had access to these models.

                                          • a-dub 3 hours ago

                                            > The multi-agent workflow led to interesting dynamics. For example, the key idea in producing this attack was discovered by a pair of workers working together. Both started investigating the idea; the first worker prematurely rejected the idea as infeasible, but the second found a way to fully exploit it. The pair kept exchanging messages, and eventually both agreed they had found an effective attack.

                                            this is pretty interesting. the way it is written doesn't make it sound like the collaboration actually led to the discovery, but rather just the stochastic nature of each thread in the search. it would be interesting to replay and repeat the search (possibly with prior/context pertubations) to get a sense for how often it finds or misses the known working path.

                                            • TeMPOraL 3 hours ago

                                              Hypothesis: the pairing / collaboration makes it much more likely to find a fruitful road previously dismissed, because... that's what happens in fiction - including books, movies, and journalism (long-form "people stories"). It's a common trope: if one character dismisses a course of action, the plot demands the other character to take it.

                                              In a way LLMs are, after all, trained to LARP people, including fictional characters and their tropes - this was actually exploited for jailbreaking to good effect in the late pre-agentic era (read: some two years ago). C.f. Waluigi effect. Not sure if it still holds for current models, but I can't imagine why it would not.

                                              • imightbebatman 3 hours ago

                                                Yes there are two interesting derivative questions from this, assuming I understand it.

                                                First, is it reproducible consistently at ~50% of workers? If not, what is the rate.

                                                Second, are there any lessons to be learned here to increase the rate of success by changing models/weights/training?

                                                The news by itself isn't really good news. But it could lead to good news. Maybe.

                                              • vuciuc 3 hours ago

                                                > But as we develop increasingly powerful cryptanalytic results, it would be prudent to consider how researchers should react if a language model were to discover vulnerabilities in cryptosystems where attacks do have an immediate real-world impact.

                                                How would they react if a human were to discover vulnerabilities in cryptosystems?

                                                • minraws 3 hours ago

                                                  First of all we likely wouldn't know it's better to call US govt or any other govt if you have that tech, and then take that govt job and hope you can happy life... instead of annoucing it publicly only when it's a AI model where we expect it's ability to tend/scale towards infinity does it become something to tell the wider public.

                                                  Although if RSA had a vulnerability I would be very very shocked probably because I still haven't learnt post quantum encryption algorithms enough to really feel like they should be unbreable...

                                                  If there is a researcher or someone in space how should I feel about it. Is it as bad as RSA being completely broken open?

                                                  I do understand that AI will get better, and a lot actually at very easily verifiable tasks but this one I find it hard to wrap my head around because of my ignorance.

                                                • ls612 3 hours ago

                                                  with black vans.

                                                  • ComplexSystems 2 hours ago

                                                    Fortunately, Claude won't fit in a van!

                                                    • dgellow 2 hours ago

                                                      Technically a claude model can fit in a pocket size hard drive

                                                • Retr0id 3 hours ago

                                                  TL;DR: They marginally improved on the best known academic attack on 7-round AES-128 (which normally uses 10 rounds - you do not need to worry about AES being broken).

                                                  The attack on HAWK is perhaps more interesting - they were able to halve the effective key length. HAWK is a candidate for NIST standardisation. It has been studied academically, but isn't really deployed anywhere (because it hasn't been standardised!)

                                                  • adrian_b 57 minutes ago

                                                    It should be noted that the attack is not only an attack against a weakened AES, but it is also a chosen-plaintext attack.

                                                    It is standard in cryptography to analyze ciphers under this kind of attack, which is stronger than normal attacks, because a cipher that resists to a stronger attack will also resist to weaker attacks, so using the strongest possible attack increases the confidence in a cipher.

                                                    While using the strongest attack for testing a cipher remains the correct method, chosen-plaintext attacks are no longer realistic today, so even when a cipher appears somewhat vulnerable to such attacks that does not imply that it is vulnerable in normal use.

                                                    The reason is that the modes of operation for ciphers where the base cipher can be attacked with chosen plaintexts are obsolete. The most frequently used modes of operation are now modes like the counter mode (e.g. in AES GCM), where it is impossible to perform a chosen plaintext attack (i.e. where you must trick the victim to encrypt a text that you choose, but in counter mode the cipher only encrypts a sequence of numbers chosen by the intended victim, which cannot be influenced by the attacker).

                                                    • baxtr 2 hours ago

                                                      How is this not the top comment?

                                                    • Stevvo 3 hours ago

                                                      Interesting they are still using "Mythos Preview" instead of "Mythos 5"; I had read from others who had access to both that Mythos 5 is less capable.

                                                      • TeMPOraL 3 hours ago

                                                        Going to guess it's more available or less overconstrained. See e.g. Fable, which is much better than Opus 4.8 and possibly than Opus 5... in the rare case of a task it doesn't punt on because of its safety guardrails.

                                                      • Diogenesian 2 hours ago

                                                        This seems like a bit of an overstatement:

                                                          Despite HAWK having survived two rounds of expert human review over a period of two years, Mythos was able to improve the best-known attack on it in just 60 hours of work—effectively cutting its key strength in half.
                                                        
                                                        since, later:

                                                          Mythos’s attack works by finding a specific, previously unexploited symmetry called a nontrivial automorphism in the lattice used by HAWK. Prior work proved that efficiently finding such an automorphism would permit an attack, but did not answer if such an automorphism was accessible in the lattice used by HAWK. The automorphism discovered by Mythos allows a faster enumeration attack that, while still exponential, means that one needs to double the size of HAWK keys to achieve the same level of security.
                                                        
                                                        Not downplaying Mythos's contribution here[1], but that first paragraph strongly hinted (at least to me) that there were no known weaknesses. "Discovering a weakness that had previously been only theoretical" is vastly different from "discovering an unknown weakness." Again: very cool Mythos was able to do this. It just seems like another case of "LLMs are good at finding concrete mathematical (counter)examples" - which is also cool! But the PR here is cynical.

                                                        ...and it is kind of incredible to think that they spent $100,000 over 3 days looking for an automorphism. Not the possibility of an automorphism, that was already known. Man.

                                                        [1] ... or focusing too hard on the strange use of mathematical language...

                                                        • recitedropper 1 hour ago

                                                          Yes, this discovery is surprisingly similar to the recent counterexamples LLMs have been finding for mathematical conjectures: a semi-novel construction, built on previous work, that feels like it was found with enormous search and an okay heuristic.

                                                          I feel like there is a pattern emerging regarding the type of novel discoveries LLMs are good at finding, but it will take some more data points to see if the trend solidifies.

                                                        • wslh 32 minutes ago

                                                          I'm looking forward to seeing similar work on SHA-256. It would be fascinating if AI could discover previously unknown weaknesses in reduced-round variants.

                                                          It would also be interesting whether AI could discover new algorithmic optimizations for SHA-256 similar in spirit to AsicBoost[1].

                                                          [1] https://arxiv.org/pdf/1604.00575

                                                          • quotemstr 3 hours ago

                                                            One attack weakens HAWK, a post-quantum cryptography cipher candidate. I don't trust these PQC things one bit. I'll use them in combination with a strong clasically-resistant cipher (in so-called hybrid encryption modes), but not alone.

                                                            There's a push to turn off the classical modes and rely entirely on PQC for both quantum and classical security. Uh... no, thank you? Why would we want to do that at this point? The classical cipher component isn't hurting anything. Awfully creepy to pushing reliance on the new thing alone.

                                                            ... especially now that we have LLM-discovered attacks on the new things.

                                                            • JuniperMesos 2 hours ago

                                                              The classical cipher component is additonal complexity in the protocol and maybe some meaningful amount of additonal time to compute and key data to store/transmit, is it not? I can see why we'd like to avoid effectively encrypting the same data twice with different protocols, one of which is known to be vulnerable to quantum computer based attacks.

                                                              • vrighter 2 hours ago

                                                                good thing quantum computers that can factor numbers have never been built. No number was ever really factored without cheating, the actual shor's algorithm has never been implemented. And we're not really any closer to

                                                                • ameliaquining 2 hours ago

                                                                  That last sentence is not true; we have gotten much closer to building a quantum computer that can run Shor's algorithm. Organizations like Google and Cloudflare have declared a 2029 deadline to completely stop depending on the security of pre-quantum algorithms; hitting that deadline is going to cost a lot of engineering resources, but they're paying that cost because they think there's too great a chance that nation-state adversaries will have scalable quantum computers by then. See https://words.filippo.io/crqc-timeline/ and the various posts linked therein, including from the aforementioned companies.

                                                                • quotemstr 1 hour ago

                                                                  Classic ciphers are damn fast and small compared to PQC. If you're doing PQC anyway, doing classical cryptography at the same time has negligible cost.

                                                                  That makes attempts to push PQC-only modes super suspicious to me. Smells like Dual_EC_DRBG.

                                                                • some_furry 50 minutes ago

                                                                  > One attack weakens HAWK, a post-quantum cryptography cipher candidate. I don't trust these PQC things one bit. I'll use them in combination with a strong clasically-resistant cipher (in so-called hybrid encryption modes), but not alone.

                                                                  HAWK is a signature algorithm, not encryption.

                                                                • Johnny_Bonk 2 hours ago

                                                                  Great now can you make opus 5 work please