11 comments

  • mrieck 1 hour ago

    This metric seems to be asking if years of research by specialists could be emulated by a few LLM api calls. Reminds me of this meme:

    https://imgflip.com/i/b366vm

    • aleph_minus_one 4 minutes ago

      The difference to this meme is:

      Often people who are critical of whether AI can lead to research breakthroughs are experts in the respective area, who nevertheless are afraid of their future career prospects in academia (getting a permanent position in academia is hard and it is deeply political who gets such a position).

      This people are thus not scared by AI per se (it's basically their daily job to devise innovations that advance their field), but their fears are that

      - because of the hype around AI the research into which they invested years, often decades will be considered "unimportant",

      - incompotent people in decision-making position will think researchers can be replaced by AI.

      • bartleeanderson 7 minutes ago

        Love that. I myself first thought. How specialized is this and at what level of human complexity is this team working.

      • Eridrus 1 hour ago

        I think this sort of small scale research on this problem is inherently pointless and will be a lagging indicator of diffusion, not a leading indicator of capability.

        The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result.

        If the AI could do this task we would see this happening in places where the economic incentives let them spend millions of dollars on this problem, not on an eval like this.

        This specific form of eval where you just ask the agent to solve it with no specific scaffolding besides GPU access (e.g. nothing like AlphaEvolve, ArchPilot, etc that try to work around model shortcomings) is also going to further trail what is possible at small scale. It's good that we at least give them execution environments now, but this feels like the experiments that were worked on figuring out how to get LLMs to do native arithmetic rather than just giving them a calculator/python env.

        • curt15 10 minutes ago

          > The Navier Stokes results used millions of dollars of tokens and thousands of parallel agents to get the result.

          Not to mention a century of theoretical foundations and innovations by humans, written for human understanding.

          • hn_throwaway_99 15 minutes ago

            Regarding the Navier Stokes theorem, one of the best posts I've seen about this (and the other OpenAI math proofs) recently is Nestor Guillen's post where he introduces "the convex hull of ideas". I love this because I've had some vague ideas about the limitations of LLMs that aligned with this, but Guillen really clarified the idea and explained it with a great analogy: https://terrytao.wordpress.com/2026/09/13/happy-those-able-t...

            Basically Guillen is arguing that LLMs are great at finding results within the "convex hull" of existing literature (i.e. their training data). They can make connections across different parts of the literature where it would be impossible for a human to be an expert in all these areas. But when it comes to truly novel, original ideas, there is no proof yet that LLMs are able to go there. That's honestly a limitation I'm really rooting for because otherwise I think the future of humanity is generally fucked.

          • rmunn 20 hours ago

            Short version of the article: no, not even close.

            Practically every paragraph is negative, with sentences like "Agents made misleading claims about their work," and "A natural question is whether the agents could have improved with larger GPU budgets. Although both improved across their runs, in the case of Fable the improvements were almost entirely due to attempted cheating." and "For Sol, the answer is less clear-cut; it did make some progress, although its method was fairly incremental and had limited applicability to the coding task. This suggests that we should be pessimistic about further GPU spending," all reinforcing the fact that LLMs aren't currently capable of this.

            My own view is "No, of course not, in fact they will never be capable of achieving good results with that technique." Because that technique will end up training the LLMs on their own output and lead to the inability to distinguish reality from hallucination. If you think I'm wrong about that, I'd be interested in hearing why.

            • robrenaud 1 hour ago

              Hallucinate -> do an experiment -> see it fails, try again. Hallucinate -> do an experiment -> it works, model innovated.

              Agentic models brush up against reality, this gives a way around the hallucination problem.

              Here is a recent talk showing that hallucination and discovery are actually positively coupled. https://www.youtube.com/live/ZNlZsI9kBm4?si=nhn4ancXu7s6qtom...

              • TeMPOraL 1 hour ago

                Ground breaking thought (ML Researchers Will Hate Me For): (some) hallucinations are actually creativity.

                • emp17344 50 minutes ago

                  In practice, hallucinations have only mitigated within verifiable domains. The vast, vast majority of human activity is unverifiable, or at least significantly less verifiable than mathematics.

                • glimshe 10 hours ago

                  We don't have to prove you wrong, you have to make a case for your position. Your statement was vague and handwavy and could be countered with another vague statement such as "They will add a feature that allows the LLM to better detect hallucinations".

                  • rmunn 5 hours ago

                    I phrased it that way because I couldn't remember the term "model collapse" at the time. But that's what I meant: that making the LLMs self-train will lead to model collapse.

                    There; now the statement is far less vague and handwavy, because I'm making a specific claim that is, AFAIK, well-understood.

                    Also, you seem to have misunderstood me a little. I didn't mean "prove me wrong", I meant "If I'm missing something, please tell me about it." More of a conversational request than staking a claim in an argument. Many people at HN seem to like to take argumentative, debate-competition stances — but I usually prefer more "Hey, let's discuss this interesting idea, point out mistakes each other is making, and learn together" kind of interactions. That's what I was asking for.

                    • phoghed 1 hour ago

                      You couldn’t remember the term model collapse because people quit saying it shortly after it was invented in 2023, because it was just cope by people praying for that outcome. Internet will have LLM output, therefore models will train on that output, then kaput. And what happenend next? Models trained on model output on purpose and just keep getting better.

                      • glimshe 4 hours ago

                        Fair enough.

                        Well, I guess the answer isn't too different. I believe model collapse is a limitation of the current AI tech but maybe not the future ones. You can see humanity as a huge model that trains itself. What is novel about AI is that we built a machine with some intelligence traits that is free of biological constraints. If we can emulate the aggregate intelligence of a civilization inside a machine, it could improve itself forever but at a much faster pace.

                    • janalsncm 20 hours ago

                      A bit too pessimistic imo. I agree that AI can’t automate things end to end, but a good deal of R&D involves kicking off a training run and babysitting it.

                      If your training run dies at 1 am and you’re sleeping, you won’t find out about it until the next day. You can lose up to 18 hours of work depending on when it happens. Based on the error it might be as simple as tweaking a single hyperparameter and rebooting, which is something LLMs are usually capable of.

                      Even just that task means I can kick off multiple runs over the weekend and have confidence they’ll finish. It’s a game changer.

                      • rmunn 20 hours ago

                        I'd classify that as an entirely different category than AI self-training. What you're describing could have been done with a short script, though which parameter to tweak and how to tweak it would be difficult to automate with a non-LLM script, so the LLM's being able to parse the error message and base the tweak on the content of the error is a definite improvement to the process there.

                        But I'd classify this as LLM being used to automate a sysadmin task, rather than calling that self-training.

                        • janalsncm 19 hours ago

                          Yeah I’m not trying to argue it is AGI, but it’s not as simple as a short script. There’s some amount of debugging involved, and no amount of if-statements could cover all possible ways a script could break.

                          In a way, “recursive self improvement” just means tools helping us to create better tools. At least that’s what the words mean.

                          • cyanydeez 10 hours ago

                            the GP posits the problem with "recursive self improvement" is the poisoned context & hallucination problem. While a strong loop that has fail safes, backups, restore points (essentially, a fancy backup system), the problem isn't that we cannot create a healthy advanced wiggum loop; it's that every step of the LLM as it grows whatever knowledge is acretes, has a chance of being either poisoned (eg, it conflates two tokens as describe different things) or wholesales fabricates a method or procedure.

                            Now humans are just as bad, but they're not moving at the speed of compute so the posion and fabricates can dissolve over time, or just, as you've noticed turning on your news, get stuck in very stupid positions. So humans are clearly capable but clearly don't tend to do this either.

                            So then we dont have a real road map. The error rates, although small, acrete at exponential levels and will wash out improvements.

                            So I also had the idea that "if we just give it enough context, surely it'll be more powerful". But the error rates hit that squarely. The larger the context grows, the more likely it hasn't properly organized its knowledge to avoid overlapping facts.

                            In programming, it's worse, because a lot of the code is purposefully "DRY" and reuseable. Everye C program has a main(); is it remembering the correct main? or any of the number of same variables?

                            You can see an LLM is powerful but it's not ominipotent. It'll suffer very much when it starts hallucinations and context poisoning.

                            So, sure you can try a super ralph wiggum loop with memory, fallback safeties, etc, but you basically then need another turtle that does the same thing, and at that point, you're positing a infinite jest of ralph wiggum loops tracking each other, recursively, forever.

                        • rcxdude 11 hours ago

                          This was an interesting post on the subject that more or less agrees: 'self-improvement' is happening with things like this but there's a lot of headwind on any 'hard take-off' where capabilities grow exponentially all on their own:

                          https://www.rameznaam.com/p/471bbae4-1163-4048-944b-18f8b0bf...

                        • What will never be capable? Neural networks? Neural networks that utilize next token prediction in their training recipe?

                        • janalsncm 20 hours ago

                          In my experience R&D has basically two axes: how innovative it is, and how well we can measure the results.

                          For the quadrant of non-innovative tasks where we already have a good way to measure performance, Claude can handle this. There is very little ambiguity, and we are basically just looking to maximize some metric under a set of constraints.

                          Many business processes are not like that. They might be conceptually simple, but it isn’t that easy to say whether a system has done a good job or not. I would say that LLMs can help with this a lot but they have bad judgement because it requires talking to people.

                          And the other, perhaps more rare issue is in problems where there is data but actually modeling it to sufficient quality or fast enough is hard.

                          • wonnage 1 hour ago

                            corollary: coding is solved, bugs are not solved

                          • elendilm 5 minutes ago

                            I don't buy this whole "AI is innovative" narrative.

                            If AI is so innovative then why am I seeing a dumb retard everytime I discuss anything even remotely innovative with it.

                            It is true that it can write decent code once the domain bounds are well defined. What I have found is that it is good at scrutinizing already written code - but only with expert supervision. Even then it most often tests our patience with stupid suggestions.

                            • Charly_HW 3 hours ago

                              The only outcome I see is that AI will become so complex that we won’t be able to rely on it for R&D, because we won’t be able to measure and prove its results. Some AI collaboration and speed of work will be impossible for humans to replicate, making it neither good nor bad—just something beyond human ability.

                              • solenoid0937 1 hour ago

                                Give it a verification loop and enough compute, and AI will soon cook algorithmic R&D like it cooked mathematics. There is just no question whatsoever of this happening, it is guaranteed.

                                • jltsiren 17 minutes ago

                                  Theoretical fields like mathematics inherently suffer from a form of Goodhart's Law. The progress in mathematics you hear of is almost never the kind of progress non-mathematicians would care about.

                                  It's always easier to measure progress according to the internal metrics of the field than to evaluate the contributions to the wider understanding of the topic. In theoretical computer science (which I'm most familiar with), people have long complained about results focusing on shaving sublogarithmic factors from complexity bounds (while making the algorithm worse in practice) and about reviewers being impressed by the technical difficulty of proofs. But progress like that is easier to measure than new algorithmic ideas or conceptual understanding.

                                  But it's not all bad. The researchers chasing the metrics are almost always genuinely interested in the topics they study. Their actual contributions mostly come from the ideas they explore while trying to achieve measurable progress. And because they are not expected to produce anything of direct value (Goodhart's Law for applied researchers), they can explore a wider range of ideas.

                                  I personally noticed this when I moved from theoretical computer science to algorithmic bioinformatics. When I start a new project, the expectation is that researchers in genomics should be using sofware that uses the new algorithms five year from now. That expectation is useful, but it's also a strict constraint on what I can afford to try.

                                  • dwroberts 1 hour ago

                                    > it cooked mathematics

                                    Weird, these are all still here? https://en.wikipedia.org/wiki/List_of_unsolved_problems_in_m...

                                    • gjm11 52 minutes ago

                                      Several of the items on that list are claimed to have been resolved by the recent OpenAI theorem-dump. Another is the Navier-Stokes question that's one of the Millennium Prize problems, also claimed to have been resolved by OpenAI's models.

                                      So: unless all those claims by OpenAI turn out to have been mistakes[1]: no, actually, those are not all still there.

                                      [1] It's certainly possible that some will. They've already retracted a few things.

                                  • thesmtsolver2 1 hour ago

                                    It will cook only as long as brute force through search space is cheap and economically viable.

                                    • gus_massa 57 minutes ago

                                      Have you seen my draft? For easy task I may write the straighforward solution, but for complicated stuff it's a mix of throwing stuff to the wall and see what sticks. It's informed search by experience, but AI also use weights to pick the attempts.

                                      • solenoid0937 41 minutes ago

                                        AI is not a calculator. OpenAI did not solve hundreds of open problems in math by "brute forcing the search space," there was a lot of intelligent and novel thought (gasp!) in the way AI chose the paths it did.

                                        Yes, AI can explore thousands of ideas at once, no, that's not "brute forcing the search space" because the search space is way larger than you think it is.

                                        • simoncion 1 hour ago

                                          > ...and economically viable.

                                          And there are so many indicators that it's not currently economically viable and will only be if retail price is massively increased and strong regulatory barriers to competitors are erected. ;)

                                          • solenoid0937 40 minutes ago

                                            Very interesting if it's not currently economically viable. Do have any sources on OpenAI's margins?

                                      • thoughtpeddler 20 hours ago

                                        How much of this can change if subsequent training runs produce models that are much better at abduction?

                                        • simianwords 11 hours ago

                                          Could it be that the companies have nerfed the models on these domains? It is a very hard thing to do because it can hurt related domains. But its not beyond the ideology of Dario - he tried it publicly .

                                          • charcircuit 20 hours ago

                                            I think the more interesting thing is was it unable to do it even knowing the solution? I feel like only testing a single innovation is biasing the current state of automated AI R&D.

                                            • Oarch 41 minutes ago

                                              If AIs start getting truly good at innovation, we may cease to recognise our reality.

                                              I saw one mathematician describe the recent OpenAI math-dump as 'alien-like' math.

                                              Imagine a world of countless new aircraft designs, fuel sources, musical genres, architectural styles. It could be bewildering.