If I sat down to read just the Dutch East India Company pages myself, at two minutes a page, eight hours a day, five days a week, it would take me about 70 years. And that’s before the newspapers. My homebrew AI lab got through the entire archive in a single twelve-hour overnight run.
Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...
You know you don't have to comment on everything you see on the internet, right? You're allowed to just keep scrolling when something isn't for you. I wonder how much joy you have in your life, I suspect very little, if anything.
I didn't mean to upset you, piratebroadcast. Like so many of us, you scratched a technological itch, and that is fine.
However, now that AI tools have supercharged us, these scratches are easier to scratch, and their outcomes, made public, flood the aether.
The aether is 99% garbage thanks to entrepreneurs big and small, so anyone scratching any kind of real itch and posting it online improves SNR, ever so slightly.
Specifically that these rabbit holes are useful to bring people to all sorts of new discoveries and skills.
It is important to point out that that's a real risk with such AI use.
Of course, it is also true that it would likely not have happened at all otherwise. Both things can be true at the same time.
__
Oh I just realized that you're the actual author and this is not the only super-thin-skinned comment.
Possibly, yeah, but the reactions here do not look like the learnings have been long-term-stable building blocks, tbh.
But anyway.
I am curious what else will show back up again when other people decide to hook up a clanker to one of the many piles of historic (and current) data we do not have the manpower to process for.
>To make this kind of research accessible, I’m open-sourcing the workflow I created for this investigation as a small toolkit, Antiquity, enabling anyone with a question and a coding agent to conduct similar historical archival investigations.
There are also VOC archives at Cape Town, also in Kew (search for the letters of Loot) which were literally looted by privateers. All these are written in High Dutch some in German. How reliable are the translations?
IMHO: The rotating rhino, meteor impact, and animated flowchart is totally unnecessary cruft that makes it look almost satirical. If this keeps up, in time, this "AAA effects" stuff is going to look like the 90s "under construction" banner gifs.
The effects are comically bad. I see the inspiration in scrolling effects that the New York Times put together, but the NYT was never dumb enough to obscure the copy text. Form follows function, and the function of a web page is to be read, not to obscure what is to be read with some stupid effect that's supposed to remind one (I suppose) of a volcano's cloud obscuring one's vision. At least "under construction" banners didn't obtrude upon the copy text.
I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.
Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.
Please just make sure to keep the ethical implications of any such work in mind.
I do not know what exactly it is you're building, but the shape also fits "weapon", and weapons do not really care about the good intentions of their author.
>Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context. Very true in my case on similar problems, my major issue was OCR relics. Reasonable mispelled words say by an uneducated person, are not that much of an issue. For the OP VOC work most letters were written by educated scribes and less of a problem. Anything before 1650 had very
different calligraphy though.
That was a fascinating read, I really enjoyed it. Literally like exploring lost knowledge. Great work and a great write-up.
I also liked the aesthetics of it and the little effects (meteorite and volcano, but please fix the rhino and the text flowing around it while it rotates).
I wonder what else could be found in such archives. Some ideas:
- Locations or routes of sunken ships and their missing cargo?
- Some pirate stories, maybe about a now-forgotten but once-legendary pirate captain?
- Unusual weather events, like snow in the summer?
I love this, one major friction I’ve seen to human coordination and advancement has been the journals in different languages
Many people don’t notice, but even Wikipedia has no normalization between articles in different languages. The language button there acts like its showing you a translated version of the article but its actually a completely different Encyclopedia and community of editors with no cross reference to the other language’s article and references at all. Articles that are stubs on the English page may be massive fully fleshed out articles in another language, and nothing native to the site or anything I’ve seen will tell you that there is more information in one variant
LLM’s can find the word associations and compare them in all languages, even if it itself doesn't innately know language
and there would be so much low hanging fruit here like this engineer found
I honestly primarily enjoy my browser performance not tanking on this $5000 workstation and actually bailed out while scrolling because it wasn't really usable.
If I sat down to read just the Dutch East India Company pages myself, at two minutes a page, eight hours a day, five days a week, it would take me about 70 years. And that’s before the newspapers. My homebrew AI lab got through the entire archive in a single twelve-hour overnight run.
Makes me wonder how much the author himself learned about the Dutch East India Company. I suspect very little, if anything. Something about these exercises reminds me of junk food: empty calories and all that...
You know you don't have to comment on everything you see on the internet, right? You're allowed to just keep scrolling when something isn't for you. I wonder how much joy you have in your life, I suspect very little, if anything.
I didn't mean to upset you, piratebroadcast. Like so many of us, you scratched a technological itch, and that is fine. However, now that AI tools have supercharged us, these scratches are easier to scratch, and their outcomes, made public, flood the aether.
The aether is 99% garbage thanks to entrepreneurs big and small, so anyone scratching any kind of real itch and posting it online improves SNR, ever so slightly.
You're missing the actual point sorokod has.
Specifically that these rabbit holes are useful to bring people to all sorts of new discoveries and skills.
It is important to point out that that's a real risk with such AI use. Of course, it is also true that it would likely not have happened at all otherwise. Both things can be true at the same time.
__
Oh I just realized that you're the actual author and this is not the only super-thin-skinned comment.
Man. Why do be like this.
I think the author learned a lot in the process. It exactly served the purpose of diving down rabbit holes.
Possibly, yeah, but the reactions here do not look like the learnings have been long-term-stable building blocks, tbh.
But anyway.
I am curious what else will show back up again when other people decide to hook up a clanker to one of the many piles of historic (and current) data we do not have the manpower to process for.
If OP followed the ideas of sorokod, there would have been no learnings at all.
If you know what you're looking for in it, you're not obliged to read a book in order either.
>To make this kind of research accessible, I’m open-sourcing the workflow I created for this investigation as a small toolkit, Antiquity, enabling anyone with a question and a coding agent to conduct similar historical archival investigations.
https://github.com/jessewaites/antiquity
There are also VOC archives at Cape Town, also in Kew (search for the letters of Loot) which were literally looted by privateers. All these are written in High Dutch some in German. How reliable are the translations?
IMHO: The rotating rhino, meteor impact, and animated flowchart is totally unnecessary cruft that makes it look almost satirical. If this keeps up, in time, this "AAA effects" stuff is going to look like the 90s "under construction" banner gifs.
What's wrong with the 90s, we had great fun.
It was fun - until it was cringey, then it became fun again due to retro-fandom
Just calling the progression while we're on it.
The effects are comically bad. I see the inspiration in scrolling effects that the New York Times put together, but the NYT was never dumb enough to obscure the copy text. Form follows function, and the function of a web page is to be read, not to obscure what is to be read with some stupid effect that's supposed to remind one (I suppose) of a volcano's cloud obscuring one's vision. At least "under construction" banners didn't obtrude upon the copy text.
Maybe. But for now, for me, I find them amusing.
I mostly agree, I still read and enjoyed it but part of me kept snagging on the fx and wishing it was way way turned down.
I loved it and it brought me joy.
Except for the rhino, I must admit I liked the effects.
Very cool.
I'm working on a similar project for contemporary political opinion media. Every podcast, blog, oped, or show cut into little pieces with the structure, speaker, quotes and nouns pulled out and cross-referenced. I bring it up because I wonder if this kind of heavy-weight preprocessing is worth bringing to historical documents as well. It would be much more expensive, initially, but afterwards allows questions get answered even cheaper than they are in your current system. It may be worth collecting interested parties and co-investing in the structured parsing.
Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context.
Please just make sure to keep the ethical implications of any such work in mind.
I do not know what exactly it is you're building, but the shape also fits "weapon", and weapons do not really care about the good intentions of their author.
>Also modern transcription and historical document scanning have a similar shaped problem - dealing with misspelled words and trying to infer their corrections from context. Very true in my case on similar problems, my major issue was OCR relics. Reasonable mispelled words say by an uneducated person, are not that much of an issue. For the OP VOC work most letters were written by educated scribes and less of a problem. Anything before 1650 had very different calligraphy though.
Recent and related (by HN's own https://news.ycombinator.com/user?id=benbreen!)
Using Opus 5.5 to discover a new eyewitness record of the dodo - https://news.ycombinator.com/item?id=49926917 - Oct 2026 (79 comments)
That was a fascinating read, I really enjoyed it. Literally like exploring lost knowledge. Great work and a great write-up.
I also liked the aesthetics of it and the little effects (meteorite and volcano, but please fix the rhino and the text flowing around it while it rotates).
I wonder what else could be found in such archives. Some ideas: - Locations or routes of sunken ships and their missing cargo?
- Some pirate stories, maybe about a now-forgotten but once-legendary pirate captain?
- Unusual weather events, like snow in the summer?
(edit: formatting)
I love this, one major friction I’ve seen to human coordination and advancement has been the journals in different languages
Many people don’t notice, but even Wikipedia has no normalization between articles in different languages. The language button there acts like its showing you a translated version of the article but its actually a completely different Encyclopedia and community of editors with no cross reference to the other language’s article and references at all. Articles that are stubs on the English page may be massive fully fleshed out articles in another language, and nothing native to the site or anything I’ve seen will tell you that there is more information in one variant
LLM’s can find the word associations and compare them in all languages, even if it itself doesn't innately know language
and there would be so much low hanging fruit here like this engineer found
Okay, for those of you that do not enjoy fun, I just added a button at the top of the page to turn off the special effects.
Scrolling on mobile is very slow, even with the effects turned off.
I liked them! Can't please everyone though.
I honestly primarily enjoy my browser performance not tanking on this $5000 workstation and actually bailed out while scrolling because it wasn't really usable.
I bet it works great in chrome tho