9 comments

  • whythismatters 8 hours ago

    >To make this kind of research accessible, I’m open-sourcing the workflow I created for this investigation as a small toolkit, Antiquity, enabling anyone with a question and a coding agent to conduct similar historical archival investigations.

    https://github.com/jessewaites/antiquity

    • yannis 7 hours ago

      There are also VOC archives at Cape Town, also in Kew (search for the letters of Loot) which were literally looted by privateers. All these are written in High Dutch some in German. How reliable are the translations?

    • sorokod 7 hours ago

      If I sat down to read just the Dutch East India Company pages myself, at two minutes a page, eight hours a day, five days a week, it would take me about 70 years. And that’s before the newspapers. My homebrew AI lab got through the entire archive in a single twelve-hour overnight run.

      Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...

      • carsoon 34 minutes ago

        One usecase I could think of is weather patterns. Weather is very hard to predict so any info from the past could help us create better models. It would be incredibly tedious to collect that by hand from millions of documents but AI could do handle this quite well. A mix of text embeddings, and llm analysis could result in a data set for weather patterns in remote/ historic areas over time. Even if was as simple as a diary that said "today it rained a lot"

        Wether there is enough info to reconstruct usable data is unknown to me but it feels like a good experiment someone could try.

        • vidarh 5 hours ago

          Surely more than if they read nothing at all because most of it would be tedious drudgery of minimal value.

          • palmotea 4 hours ago

            >> Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...

            > Surely more than if they read nothing at all because most of it would be tedious drudgery of minimal value.

            It's an archive. It's all "tedious drudgery" until you figure out the value.

            Sometimes figuring out the value means reading things until notice something, which could be a pattern or something dispersed.

          • dyauspitr 6 hours ago

            Probably a lot. I’ve never learned about so many disparate subjects as I have over the last two years with LLMs.

            Why is it always these supremely weak arguments and rationalizations against LLMs that come from people that have been intelligent, at least based on their comment histories, for so many years. It’s radicalizing me. I want a data center everywhere and I want tokens to be so cheap they’re like electricity or water.

            • palmotea 4 hours ago

              > Why is it always these supremely weak arguments and rationalizations against LLMs that come from people that have been intelligent, at least based on their comment histories, for so many years. It’s radicalizing me. I want a data center everywhere and I want tokens to be so cheap they’re like electricity or water.

              What do you think the impact of "tokens to be so cheap they’re like electricity or water" will be? I think people who are intelligent and don't have their heads stuck in the sand see the implications of that. Why should people with money pay people to for intelligence, when they can buy machine intelligence very cheaply?

              I guess it will free up smart people to finally take jobs that don't use their intelligence, or sit around scraping by on minimum-wage-like UBI (if we're so lucky).

              • howunfortunate 2 hours ago

                > sit around scraping by on minimum-wage-like UBI (if we're so lucky).

                If every cognitive task currently done by humans could be done for ~free, our global material abundance would be truly unprecedented, nigh unlimited.

                I have many worries about a world where people aren't needed, but "scraping by" does not describe that world in any sense.

                • fv3y 25 minutes ago

                  The big problem with this argument is that the current owners of frontier models are actively working against these ideals.

                  Many people, including myself, would feel very differently if models were open(and not just open weights), compute/resources were cheap enough to enable local access and ownership.

                  Unfortunately, despite what they may say in press releases, a lot of people in the space are banking on holding a stranglehold over the market and building a regulatory and resource moat to protect their investment.

                  Of course there is the argument that they deserve some return on the investment in training but lets not forget that it is our work, the work of the commons, that enabled it in the first place. People are frustrated that a small group of people endeavour to control the use of a tool that was created using the work, thoughts and content of us all, often using questionably legal means.

              • tyromaniac 5 hours ago

                Soon they might be as cheap as electricity, even if token price doesn't change

                • taneq 2 hours ago

                  Ah, the ol’ switcheroo!

                • customguy 3 hours ago

                  Learning "about" a subject and actually learning a subject are totally different things.

                  • dyauspitr 2 hours ago

                    It’s just a matter of what you focus on don’t kid yourself. These things are reading more papers than you can ever read compiling them and then giving you that information along with sources. It makes learning everything better.

                  • sandworm101 5 hours ago

                    At school i was made to read all of shakespeare. It wasnt about passing tests. Any LLM can read shakespeare and pass a test. I was made to read shakespeare so that i would appreciate language in the hope that i would strive to improve my own. That still counts.

                    I am a soldier and language is vital in every day of my job. Tone, word choice, cadence ... all of it conveys meaning in a way that an LLM simply cannot. Soldiers will not follow an AI up a hill, nor will they follow those they know use AI to pretend they understand a subject.

                    Want to sound smart? Watch blackadder. Want to win no-win aguements through wit? Watch Archer. Want to inspire your troops to do somthing unpleasant? read and watch shakespeare. An LLM can teach you nothing that really counts.

                    • vidarh 5 hours ago

                      None of this is relevant to the argument above. There is no reason to believe these archives are equivalent to Shakespeare even if one were to agree with you that Shakespeare is that important to read.

                      > An LLM can teach you nothing that really counts.

                      Nothing you wrote supports that argument.

                      None of us can read everything. None of us can even read every novel published in a single month. So we filter. One way an LLM can teach you is as an excellent way of filtering information and give you a chance to read what really counts.

                      • sandworm101 4 hours ago

                        Ya. Try reading a few hundred pages of historical documents. Expertise does not come from the facts that an LLM can so easily integrate. It comes from understanding the mindset of the people who wrote those pages. Do it properly and you should start to speak and write like those people. The path to becoming an expert on the East India cannot be shortcuted via an LLM.

                        • zozbot234 2 hours ago

                          LLMs actually do a very fine job of understanding the "mindset" embodied in historical sources, because historical documents themselves are a valuable source of highly diverse LLM training data. You can easily point pretty much any model to any random historical text, no matter how obscure (probably not an archival log entry, though) and ask "what is this text about in modern language, how might it be of interest or relevant today and what points would be considered especially outdated?" you will generally get very helpful answers. Even the thinking output is instructive.

                  • scotty79 5 hours ago

                    I'm sure that if he didn't do what he did, he would lear about Dutch East India Company so so very much.

                    Remember that even junk food is more food than junk and you can survive on it for years.

                  • jttnr 6 hours ago

                    That was a fascinating read, I really enjoyed it. Literally like exploring lost knowledge. Great work and a great write-up.

                    I also liked the aesthetics of it and the little effects (meteorite and volcano, but please fix the rhino and the text flowing around it while it rotates).

                    I wonder what else could be found in such archives. Some ideas: - Locations or routes of sunken ships and their missing cargo?

                    - Some pirate stories, maybe about a now-forgotten but once-legendary pirate captain?

                    - Unusual weather events, like snow in the summer?

                    (edit: formatting)

                    • jvanderbot 7 hours ago

                      IMHO: The rotating rhino, meteor impact, and animated flowchart is totally unnecessary cruft that makes it look almost satirical. If this keeps up, in time, this "AAA effects" stuff is going to look like the 90s "under construction" banner gifs.

                      • Telemakhos 7 hours ago

                        The effects are comically bad. I see the inspiration in scrolling effects that the New York Times put together, but the NYT was never dumb enough to obscure the copy text. Form follows function, and the function of a web page is to be read, not to obscure what is to be read with some stupid effect that's supposed to remind one (I suppose) of a volcano's cloud obscuring one's vision. At least "under construction" banners didn't obtrude upon the copy text.

                        • wavewrangler 1 hour ago

                          I quite liked the effects...I found that they didn't actually impede my reading because it only covers a small, moving portion of the text at any given time. Also, I think this was written for a more general audience, including kids. This isn't presented as a journal submission so I think the added flare is fine, personally

                        • yannis 7 hours ago

                          What's wrong with the 90s, we had great fun.

                          • jvanderbot 7 hours ago

                            It was fun - until it was cringey, then it became fun again due to retro-fandom

                            Just calling the progression while we're on it.

                          • qarl 7 hours ago

                            Maybe. But for now, for me, I find them amusing.

                            • zamadatix 4 hours ago

                              I say we bring back this kind of whimsy like we had with early GeoCities where things could be "bad" but fun instead of "quality" but formal.

                              • xenospn 7 hours ago

                                I loved it and it brought me joy.

                                • thenthenthen 4 hours ago

                                  Also all the fonts, its quite unreadable in general somehow even before the 'effects'

                                  • acgourley 7 hours ago

                                    I mostly agree, I still read and enjoyed it but part of me kept snagging on the fx and wishing it was way way turned down.

                                    • scotty79 5 hours ago

                                      I didn't read it beyond few fragments about meteorite. But I enjoyed the rhino.

                                      • jttnr 6 hours ago

                                        Except for the rhino, I must admit I liked the effects.

                                        • IAmGraydon 4 hours ago

                                          The distance between an idea and the manifestation of said idea is now nearly zero. That goes for both the good and bad varieties. That’s pretty cool, but also pretty terrible.

                                        • acgourley 7 hours ago

                                          Very cool.

                                          I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.

                                          Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.

                                          • yannis 7 hours ago

                                            >Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context. Very true in my case on similar problems, my major issue was OCR relics. Reasonable mispelled words say by an uneducated person, are not that much of an issue. For the OP VOC work most letters were written by educated scribes and less of a problem. Anything before 1650 had very different calligraphy though.

                                            • hypfer 7 hours ago

                                              Please just make sure to keep the ethical implications of any such work in mind.

                                              I do not know what exactly it is you're building, but the shape also fits "weapon", and weapons do not really care about the good intentions of their author.

                                              • acgourley 5 hours ago

                                                I hear you, I think on balance it's good which is why I'm working on it. It makes elite opinion legible and helps detect organized disinformation dark matter. It's not like the intelligence and advertising markets needs help from me about how to surveil downwards.

                                            • dang 7 hours ago

                                              Recent and related (by HN's own https://news.ycombinator.com/user?id=benbreen!)

                                              Using Opus 5.5 to discover a new eyewitness record of the dodo - https://news.ycombinator.com/item?id=49926917 - Oct 2026 (79 comments)

                                              • wavewrangler 2 hours ago

                                                In my experience, when a disruptive tech comes along that displaces a certain way of existing, all of the former examples of this happening have resulted in people adapting. That's what must happen. And if you think LLM's are lessening the value of previously valuable work, then it is time that you increased your own output to once again be high and above that which LLM's are replacing or to what you perceive as having lost value. You can do that, that is totally within you, but you have to find it for yourself. Or you can just continue talking smack and contributing to nothing. But then your output is going the opposite direction of what you say LLM's are taking away to begin with. These are conflicted times we live in, but they really don't have to be.

                                                For the record, I think this project was an excellent presentation, I loved the interactive elements snd the meteorite (or is that a meteorwrong?) to come across the page. This is about as good of a use of LLM's as I have seen. I didn't see anything about it in the article, but does anyone know what future plans are for this particular project, or is that a wrap? I didn't see I the Where This Stands section anything about future search topics. This really feels like a time where finding the question is every bit as important as finding an answer to that question.

                                                • fudgybiscuits 4 hours ago

                                                  Sorry he found an unrecorded volcanic eruption in some records? He found "A new eyewitness record of the extinct dodo" that's 400 years old? These seem to me to be rather dubious claims.

                                                  • IAmGraydon 4 hours ago

                                                    The final dodos went extinct only about 400 years ago, so not that surprising.

                                                  • yieldcrv 6 hours ago

                                                    I love this, one major friction I’ve seen to human coordination and advancement has been the journals in different languages

                                                    Many people don’t notice, but even Wikipedia has no normalization between articles in different languages. The language button there acts like its showing you a translated version of the article but its actually a completely different Encyclopedia and community of editors with no cross reference to the other language’s article and references at all. Articles that are stubs on the English page may be massive fully fleshed out articles in another language, and nothing native to the site or anything I’ve seen will tell you that there is more information in one variant

                                                    LLM’s can find the word associations and compare them in all languages, even if it itself doesn't innately know language

                                                    and there would be so much low hanging fruit here like this engineer found

                                                    • zozbot234 2 hours ago

                                                      > Many people don’t notice, but even Wikipedia has no normalization between articles in different languages. The language button there acts like its showing you a translated version of the article but its actually a completely different Encyclopedia and community of editors with no cross reference to the other language’s article and references at all.

                                                      The Wikipedia folks are working on language-independent, machine-readable structured representation of encyclopedic text (probably relying on something very much like frame semantics, via some sort of general compositional structure) in order to address this issue - see Abstract Wikipedia but note that the project is still at a very early, highly experimental stage. LLMs are not considered adequate for this task because they are non-deterministic and not auditable by humans, hence why a different approach is being planned.

                                                      • yieldcrv 1 hour ago

                                                        LLMs would be good enough until this project got anywhere

                                                        Don’t let perfect get in the way of good

                                                        I don’t trust Wikipedia to be incentivized to do this, and looking at the stub about it I have even less confidence

                                                        https://en.wikipedia.org/wiki/Abstract_Wikipedia

                                                        the last reference is from 2023 before this evolutionary branch of transformers at all