Yep, it's Murphy's law of online content. Anything you want to reference later will be gone, with no archive copies. Anything you want deleted will be available forever.
Am I getting old? 09-14 is not even close to the old web for me. The old web, to me, was back when people still published physical 'phone' books for websites.
I vaguely recall those, but they were more for normies trying to get online. For me the old web is what I saw when I logged into my university gopher server and saw the advertisement for something called the World Wide Web which I could browse via lynx. Soon enough I got a PPP connection and then Mosaic/Netscape 1.0. However everything after javascript shipped (let alone CSS) is new new new. I'd almost go as far as saying if it doesn't have a tilde in the URL it's not old web... almost...
The old web doesn't necessarily mean the oldest web. 12-17 years ago was very much an older fairly different era of the web that's worth analyzing even if it's on the younger side of the old web. I can definitively sympathize with your reaction though, it doesn't feel like that era was that long ago yet.
The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration. Not counting cost for tokens to do the development.
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
The ENTIRE thing is AI generated. I'm not talking about the article. I'm talking about the entire website, the entire "product". https://0.mk/blog/zero-humans
Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.
> The name was registered in 2009 because it was the shortest URL possible: a zero, a dot, two letters. Seventeen years later the zero means something else.
Did the zero ever mean anything? It's still a 3 character domain regardless if the first character is a zero.
Webpages dying is probably one of the biggest design flaws of the original web.
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.
We hot-linked to all those image hosts because we couldn't imagine them disappearing.
Archive.org had incredible foresight and if it didn't already exist, I'd call such a project a pipe dream.
Even Archive.org is rather limited in what it keeps. I know of a very large site that recently disappeared. archive only has part of the web html part of the site. Everything else is either gone or non accessible.
...and I had fun replacing images people directly linked from my server with less-- ahem-- savory images.
I enjoyed the emails I got from a couple people who were adamant I "hacked" their site because their "web developer" linked to images on my server... images that now said stuff like "I'm a loser bandwidth thief!", etc. (I never did use really nasty "shock" images-- just taunting stuff.)
The technology is still in its infancy unfortunately, so there's no way the web could have been based on it, but I think content-addressing is the long-term play. If I click a link, there are some cases where I want the server to respond with a fresh response just for me (e.g. a website showing the weather). But often I just want whatever content was linked to (e.g. a webpage explaining a math content). In the latter case, it would be nice if the link had a hash of the content in it, and 3rd parties could host copies to keep the link working even if the original operator stopped existing.
It's pretty difficult to avoid without very significant tradeoffs, though. The closest is content-addressable peer-to-peer networks, but these still rely on someone keeping the information around, and they struggle to scale anywhere near as much.
I found an old database backup of 0.mk on a disk I had kept.
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded.
- The first link ever shortened was a CSS stylesheet on a WordPress blog.
- Someone shortened localhost on the second day.
- The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
2009-2014? That's not the old web. Or I'm old. Take your pick.
Remember the good ol' days when we all thought that everything on the web would exist for eternity and over.
No, that's just embarrassing stuff. That stays on the web forever.
Yep, it's Murphy's law of online content. Anything you want to reference later will be gone, with no archive copies. Anything you want deleted will be available forever.
Purevolume mention made me sad. I miss that community.
Am I getting old? 09-14 is not even close to the old web for me. The old web, to me, was back when people still published physical 'phone' books for websites.
I vaguely recall those, but they were more for normies trying to get online. For me the old web is what I saw when I logged into my university gopher server and saw the advertisement for something called the World Wide Web which I could browse via lynx. Soon enough I got a PPP connection and then Mosaic/Netscape 1.0. However everything after javascript shipped (let alone CSS) is new new new. I'd almost go as far as saying if it doesn't have a tilde in the URL it's not old web... almost...
The old web doesn't necessarily mean the oldest web. 12-17 years ago was very much an older fairly different era of the web that's worth analyzing even if it's on the younger side of the old web. I can definitively sympathize with your reaction though, it doesn't feel like that era was that long ago yet.
Seriously. 2009 is several years after everyone was already saying “web 2.0”! That is nowhere near the “old web”.
The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration. Not counting cost for tokens to do the development.
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
100% AI generated text. The irony.
Proof? Using an AI checker isn't accurate.
The ENTIRE thing is AI generated. I'm not talking about the article. I'm talking about the entire website, the entire "product". https://0.mk/blog/zero-humans
Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.
> The name was registered in 2009 because it was the shortest URL possible: a zero, a dot, two letters. Seventeen years later the zero means something else.
Did the zero ever mean anything? It's still a 3 character domain regardless if the first character is a zero.
"The old web" is 1993 to 2007. It's all been downhill ever since.
Webpages dying is probably one of the biggest design flaws of the original web.
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.
We hot-linked to all those image hosts because we couldn't imagine them disappearing.
Archive.org had incredible foresight and if it didn't already exist, I'd call such a project a pipe dream.
Even Archive.org is rather limited in what it keeps. I know of a very large site that recently disappeared. archive only has part of the web html part of the site. Everything else is either gone or non accessible.
That's not my experience
> We hot-linked to all those image hosts because we couldn't imagine them disappearing.
No, we hot-linked all those image hosts because we didn't want to pay to host it ourselves.
...and I had fun replacing images people directly linked from my server with less-- ahem-- savory images.
I enjoyed the emails I got from a couple people who were adamant I "hacked" their site because their "web developer" linked to images on my server... images that now said stuff like "I'm a loser bandwidth thief!", etc. (I never did use really nasty "shock" images-- just taunting stuff.)
The technology is still in its infancy unfortunately, so there's no way the web could have been based on it, but I think content-addressing is the long-term play. If I click a link, there are some cases where I want the server to respond with a fresh response just for me (e.g. a website showing the weather). But often I just want whatever content was linked to (e.g. a webpage explaining a math content). In the latter case, it would be nice if the link had a hash of the content in it, and 3rd parties could host copies to keep the link working even if the original operator stopped existing.
It's pretty difficult to avoid without very significant tradeoffs, though. The closest is content-addressable peer-to-peer networks, but these still rely on someone keeping the information around, and they struggle to scale anywhere near as much.
Wow, that is a great domain!
Site got hugged? Is there a torrent?
I found an old database backup of 0.mk on a disk I had kept.
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded. - The first link ever shortened was a CSS stylesheet on a WordPress blog. - Someone shortened localhost on the second day. - The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
How did you managed to obtain that domain? Usually single digit or letter domains are “reserved”.
.mk appears to allow for it: https://marnet.mk/wp-content/uploads/2023/01/pravilnik-mk-mk...
“The name of the .mk domain consists of a minimum of 1 (one) and a maximum of 63 characters.”
They’re not reserved, pretty easy to obtain if you have money, starting from less than $1k.
ICANN might have set some rule for .com/.net/.org but it's not universal for all tlds