Xorshift Generators

(alanzucconi.com)

91 points | by tobr 4 days ago

7 comments

  • jabl 4 hours ago

    A number of years ago I implemented xoshiro256** for the GFortran compiler. Previously it used Marsaglia's KISS generator, which wasn't bad but perhaps no longer state of the art on the TESTU1 etc. tests. Additionally, xoshiro256** can be used in parallel by multiple threads; that took a bit of clever hacking to work around the limitations of the Fortran intrinsics API.

    • larsbrinkhoff 6 hours ago

      Xorshift-36 for the PDP-10, with help from Sebastiano Vigna.

      https://github.com/larsbrinkhoff/xoroshiro-36

      • delduca 4 days ago

        I’ve replaced Lua’s random by this. I’ve posted about it here https://nullonerror.org/2025/08/02/replacing-lua-s-math-rand...

      • YZF 8 hours ago
        • saithound 4 days ago

          Ah yes, Xorshift, the RANDU [1] of the 21st century [2].

          There is no real use case for better non-CS generators, as explained by adrian_b back in 2021 [3].

          [1] https://en.wikipedia.org/wiki/RANDU [2] https://arxiv.org/abs/1908.10020 [3] https://news.ycombinator.com/item?id=28886698

          • moregrist 4 days ago

            I’m not sure what point you’re trying to make, exactly, but a use case for better non-CS generators has always been stochastic simulation, especially simulation/sampling approaches that are bound by the number and quality of uniform variates per second.

            As someone who has spent considerable time working in these areas, I still appreciate advances.

            • saithound 4 days ago

              > I’m not sure what point you’re trying to make,

              Have you skimmed the linked thread?

              > especially simulation/sampling approaches that are bound by the number and quality of uniform variates per second

              Sorry, nobody does stochastic simulations where the number of uniform random numbers obtained per second is any sort of bottleneck. If you've spent considerable time on stochastic simulation, you already know this.

              But even if you insist that you alone are doing some very weird stochastic simulation which is somehow bottlenecked on sourcing random numbers fast enough, the falling in planes phenomenon linked above would make xorshift-type generators a poor choice for most sorts of simulations. It introduces spatial correlations into any sort of lattice dynamics simulation (Ising model, percolation) and every high dimensional Monte Carlo integration. Beyond falling in the planes, since xorshift is linear over GF(2), it is also a particularly bad choice for nondeterministic cellular automata and Boolean dynamical systems which use parity, bit masks, or xors.

              AES-CTR throughput on a modern CPU is higher than that of xoshiro256++, and much higher quality. No advances in non-CS PRNGs can beat that while maintaining the same quality. If your stochastic simulation is bottlenecked on random bits, CSPRNGs are still the way to go, and they don't interact in nasty ways with any dynamical system you can actually sinulate quickly.

              • moregrist 4 days ago

                > Sorry, nobody does stochastic simulations where the number of uniform random numbers obtained per second is any sort of bottleneck. If you've spent considerable time on stochastic simulation, you already know this.

                Actually, I spent a considerable amount of time in my doctorate and postdoc doing this.

                Any kind of MCMC sampling of a simple model tends to be bound by the rate you can draw variates.

                Examples of this include: Gillespie simulations of chemical kinetics, Ising and Potts lattice models (including their roughly bazillion variations), and anything resembling bootstrap or permutation sampling.

                Just because your problems aren’t bound by the rate of drawing uniform variates doesn’t mean that these problems don’t exist. It just means that you have a narrow view.

                • kbolino 3 days ago

                  Whether your recommendation is valid seems to be quite CPU-dependent. Cf:

                    cpu: AMD Ryzen 5 5600X 6-Core Processor             
                    BenchmarkAES_CBC-12             100000000               10.96 ns/op
                    BenchmarkAES_CTR-12             83161234                14.36 ns/op
                    BenchmarkPCG-12                 345336063                3.463 ns/op
                    BenchmarkChaCha8-12             174143492                6.894 ns/op
                    BenchmarkXoshiro256p-12         254343658                4.717 ns/op
                    BenchmarkXoshiro256pp-12        266837442                4.496 ns/op
                  
                  vs.

                    cpu: Apple M4 Pro
                    BenchmarkAES_CBC-14             162698020                7.370 ns/op
                    BenchmarkAES_CTR-14             242501074                4.954 ns/op
                    BenchmarkPCG-14                 197000988                6.083 ns/op
                    BenchmarkChaCha8-14             237430095                5.050 ns/op
                    BenchmarkXoshiro256p-14         252911710                4.738 ns/op
                    BenchmarkXoshiro256pp-14        252656401                4.745 ns/op
                  
                  Code: https://gist.github.com/kbolino/afbb86f3c9b2bd2f87272801d156...
                  • dgacmu 4 days ago

                    This is a very weird hill to die on.

                    I do a lot of testing and designing of things like hash tables and filters, and having a really fast, non-CS generator is incredibly useful for being able to clearly identify performance bottlenecks in designs. PCG has been spectacularly useful for that purpose for me.

                    • saithound 4 days ago

                      When was the last time a new PRNG helped you clearly identify a performance bottleneck?

                      As in, you were using state of the art generator X, and you couldn't see the performance bottleneck, but updating to a newer (faster, or same speed but higher quality) generator Y, and could subsequently identify the performance bottleneck?

                      If you're using PCG, not in the last 12 years.

                      (In a parallel comment I suggest trying AES-CTR for this use case)

                      • dgacmu 4 days ago

                        It's not critical but if you gave me something that behaved statistically like PCG (i.e., I didn't fret about whether it was going to cause me weird problems) but was twice as fast I'd be happy and would shift to it - it would speed up profiling and measuring and that would be nice. We still find ourselves often pre-generating a list into memory to keep the prng entirely off of the measurement path. It wouldn't be magic, but I don't need magic. I like nice things that make my life a little easier in a small corner of my research. :)

                        • Straw 4 days ago

                          Most modern RNGs should be faster than memory bandwidth (when optimized), so unless your list is small enough to fit in cache, its unclear if this is faster?

                        • saithound 4 days ago

                          dgacmu: if you're writing C on x64, try AES-128-CTR (AES-NI, 8 way) using the header wmmintrin.h which has hardware accelerated primitives for this. An LLM can implement the RNG for you based on this comment if you want to test it out quickly. It should be faster than PCG, and higher quality.

                          • dgacmu 4 days ago

                            Will do. I'm on vacation right now and losing my laptop for a few days, but seems worth trying. My recollection from the RNGs a decade ago (I'm dating myself) was that the AES approaches had higher latency but were quite decent, though slower than PCG. Curious how that's evolved.

                          • magicalhippo 4 days ago

                            > When was the last time a new PRNG helped you clearly identify a performance bottleneck?

                            While not a bottleneck as such, I contributed to a photorealistic path tracer using the Metropolis algorithm[1], and we got a 10-15% increase in samples/second when we switched from a decent to a much faster and better PRNG. Like you we didn't think the performance of it mattered much until we profiled it.

                            Granted this was a decade or so ago, would be interesting to compare the state of the art PRNGs.

                            Anyway, just pointing out that there can be real-world cases.

                            [1]: https://en.wikipedia.org/wiki/Metropolis_light_transport

                        • SideQuark 4 hours ago

                          As others here have pointed out, this is nonsense. The vast majority of PRNG calls on the planet are extremely high perf simulations, where crypto secure versions are a ludicrous cost in speed, energy, and sheer stupidity. That you and others do not understand is simply because you don’t see the places it’s required.

                          I’ve a PhD, have written papers on PRNGs, have worked in both cs prng and high perf prngs, have done decades of HPC projects, scientific sims. I get called in to develop precisely these high performance systems, and when you want to replace trillions to quadrillions of PRNG calls with one costing 10-1000x more, you’d get deservedly fired immediately.

                          You keep arguing about AES style code on a CPU. That’s not where people do high performance code. Try implementing AES and a fast prng on a GPU. You’ll soon find out how absolutely terrible cs-prngs are at performance. The measuremt isn’t how many ns per prng. It becomes how many thousands of prng generated per ns.

                          It’s bafflingly shortsighted for people with zero work in this area to continue to argue this. Choose the right tool for the job. Don’t project ignorance as knowledge. Both are useful advice.

                      • kevin_thibedeau 6 hours ago

                        > no real use case

                        Yes there is. Not every system has the need or the resources to maintain a secure random sequence. You may also want a reproducible pseudo-random sequence in generative code that logs the seeds. Because of the misguided attitude that nobody needs these features, everyone who does need them has to roll their own now.

                        • Dylan16807 6 hours ago

                          > Not every system has the need or the resources to maintain a secure random sequence.

                          I'm sure there's something, but that category has to be shrinking every year. What does such a system look like this decade, that needs random numbers but can't easily implement something like AES?

                          > You may also want a reproducible pseudo-random sequence in generative code that logs the seeds. Because of the misguided attitude that nobody needs these features, everyone who does need them has to roll their own now.

                          I don't know what difficulty you're referring to. Basically every CSPRNG can be seeded easily and you can log the seed.

                        • mallets 4 days ago

                          I don't know about all that, but I use Marsaglias for generating noise samples in MCUs like Pico. It's the fastest option there is for such devices.

                          • AlanZucconi 3 days ago

                            I agree! For many games (especially the ones running on old devices), xorshifts are pretty good! Not all applications need cryptographically safe generators! And sometimes it's ok to trade complexity for speed!

                            But I agree that if you're building something new aimed for modern devices, Marsaglia's xorshift128 wouldn't be my first choice! But I'd definitely want it in my RNG library for backward compatibility!

                            • WalterGillman 32 minutes ago

                              If you are interested in performant retro PRNGs, you can find a link to mine in my profile, if I am not mistaken. The counter is based on xorshift. As of the time of release it passed all tests that could be passed with only a 2^32-1 period.

                              On Z80 it was only about as fast as RC4 which is also a good option in terms of performance and quality but whereas RC4 has a huge state, this one only has a 32bit state, the rest staying in ROM.

                        • AlanZucconi 4 days ago

                          I'm really curious... How did you manage to post this link before I did?

                        • smusamashah 4 days ago

                          > No multiplications. No divisions. No lookup tables. Just a few bitwise instructions.

                          Would have appreciated this article more if it was written by a human.

                          • AlanZucconi 4 days ago

                            I understand that a big portion of online content is now 100% AI-generated, and that can be somewhat problematic. But for creators like me, who have been publishing articles and books for over 10 years, this AI witch hunt can be quite demotivating.

                            I worked on this project for over one year. I wrote an entire distributed framework to calculate maximal triplets, and I have 130+ machines running 24/7 for 12 weeks on N=8192. This article is an extended version of the script for the video documentary that will be released before the end of the year.

                            If you look back at my website, I used to publish two small articles a week. I've since reduced to 1 or 2 large pieces a year. And one of the reasons was exactly to rise above the many blogs that post small, fragmented articles, which could be generated in 2 minutes by ChatGPT. If I wanted to continue in that direction, I could be publishing 100 short articles a week with ChatGPT.

                            I started writing practical shader tutorials back in 2014 because there were not enough good, accessible resources online. I'm now focusing on large, in-depth pieces with original research, because that's what's valuable right now that ChatGPT has replaced StackOverflow.

                            If you appreciated this article, I hope you'll appreciate it even more knowing that YES, it was written by a human (it's me, hi!). <3

                            • smusamashah 3 days ago

                              Then I am very sorry about that, I feel bad and should read more carefully next. I only glanced at first few paragraphs and encountered this line which is a very common LLM trope. I could see that article is long and has lots of images and other stuff but I could see em dashes too. As no one else had said that yet, I did.

                              But I do want to say that this 'witch hunt' is OK. We should keep calling out AI content which is mostly effortless to keep actual effort by humans separate from this flood of content.

                              • AlanZucconi 3 days ago

                                I could have never imagined that my long love for em dashes would have been my downfall! :)

                                I think your sentiment is totally understandable. The internet is flooded with AI slop. But not all AI content is automatically slop.

                                We're moving into a new phase of the Internet were most (all?) content will eventually be touched by AI in some ways. Whether it's Grammarly fixing the grammar, ChatGPT coming up with content corrections, Claude helping with the code, and Copilot writing the integration tests.

                                And I don't necessarily see a problem with 100% AI content either, when it's actually good. Let's be honest: most people can't even write as eloquently as ChatGPT does! XD I think what I'm NOT ok with is someone pretending they didn't use AI when they did. But as long as the contribution is clearly labelled, I think I'd be ok with that!

                                The problem is that, like all "witch hunt"-adjacent phenomena, it inevitably gather momentum from haters and so-called "content destroyers". And this, paradoxically, will make creators (like me) less likely to make new original content.

                                • cmrx64 8 hours ago

                                  we should call out effortless hunters too, you cast more asperity with your folly than you find truth. try pangram next time, but I do want to say that you aren’t OK.

                                  • larsbrinkhoff 6 hours ago

                                    LLMs adopted phrases from humans. But now humans adopt phrases from LLMs.

                                    • Brian_K_White 3 hours ago

                                      Obviously the witch hunt is not ok.

                                      I mean you just exhibited why we even have the term "witch hunt" and why it's recognized as a bad thing in the first place. You burned a witch who was not a witch.

                                      I guess you must surely be fine with "we should keep calling out the low effort thoughtless critics attacking people who didn't harm them".

                                    • makira 4 days ago

                                      I enjoyed the article, and thank you for sharing it. Don't be demotivated by the current AI witch hunt, which I see as a form of tribal signaling that will pass.

                                    • tptacek 8 hours ago

                                      This is unbelievably good work, thank you for writing it!

                                      • SideQuark 4 hours ago

                                        Same - very good article.