Graydon Hoare, the creator of the Rust programming language, wrote his seminal piece on text in 2014. In it, he said "text is the most powerful, useful, effective communication technology ever, period."[1] Text is durable.
[1] https://archive.ph/FhG5L (the original either got deleted or login-walled, here is the archived version)
I love plain text, but the calculus of being able to access the content in 50 years is much more interesting in the age of agents. They can infer the meaning of structured data without schemas, decompressed archives, find embedded files, etc. It's not perfect, but the durability of binary formats is better now than it has ever been.
It's not baked-in to the weights of any model, but agents can write their own tools to work with arbitrary binary formats and get many of the benefits of off-the-shelf unix utilities.
I love that the Godot game engine has a textual scene description. That's a stark contrast to Unreal Engine's binary format for blueprints (visual scripting).
This is the public dividend of a standard finally winning. The article gives credit to Unicode, but it is the fact that ASCII unambiguously won that gives plain text its portability and longevity. It looks like Unicode is on its way to winning in the same way, but it is not there yet. Most text files I write are still pure ASCII because that's the only way to avoid unexpected glitches [0].
Ahh, plain text, wherein "plain" does some heavy lifting. If plain text had something as simple as a 4 byte signature it could have been soooo much better.
As it is programs have to guess the following;
A) encoding. If ANSI which code page? If unicode which encoding? In the case of utf-16 big or little endian?
B) line endings? CR? LF? CRLF? I guess no-one uses LFCR... right?
C) number formats? 100.000 or 100,000? Date formats? is that mm-dd-yyyy or dd-mm-yyyy?
D) what human-language is it in?
E) CSV? Don't get me started...
Yes. Text is a lot easier to load, parse, make guesses about than say XLS. Yes all of the above things can be "guessed" to a greater or lesser extent.
Yes BOM exists to at least try solving the encoding question. Of course most files don't use them. Lots of tools don't support them. And they don't solve any of the other issues.
But sure, ASCII files in English with US date formats, and Windows line endings....no problems at all...
What, no mention of our old friend, ASCII-armored Base64? For shame! Everything can be text with Base64, including things that have absolutely no business being text! Best of all, in light of popular widespread abuse of every available resource, Base64 is comparatively efficient! Bring on the petabytes! Yay and I'm not being completely sarcastic
ASCII (7-Bit) is the only widely understood charset there is. Everything beyond this point depends on the loaded charset.
The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.
Unicode are multibyte characters with variable byte length and endianess at play. If you read it wrong or guess the length wrong your results might be anything but useful.
Fun timing on this plaintext conversation. I built a little browser based plaintext playwriting app this weekend for a little weekend project. Found myself really hating 1. how clunk screenwriting software can be and 2. How unsharable the files are.
With plaintext you can hand the file off to anyone on any device--the caveat being absolutely no one wants to be handed a plaintext script. The software I used ten years ago to write plays is long since deprecated and those files are basically unopenable. Plaintext however remains.
John August has also talked about using plain text for movie/television screenplays, specifying a format based on markdown and hollywood standard: https://fountain.io/
Sharing aside, I imagine a plain text format would also be helpful for version control.
My understanding is that HN has started incorporating AI tooling in the back-end to scan for LLM-generated submissions and source content, in an effort to discourage it and encourage human-made content (with human discussions, one would hope). Why do we keep having these kinds of articles every day? Entire "apps" are generated - see the lighthouse one - and submitted, and are plain-as-day LLM-generated nonsense.
Where did you get that understanding from? I might have missed something. They penalize LLM comments. Do they also penalize submissions? Any evidence of that?
The recently released models have improved their writing styles. I don't know about this article, but if that's any indication of their writing habits, the author's github pages repo is full of posts generated by Claude.
The whole structure had a faint smell to it. It's hard to put into word, there was no obvious "it's not the quiet part. It's the load bearing bedrock", but the way the thing flowed made me strongly suspect LLM involvement.
perused the posts and incidentally all are touching high-octane topics, posted only recently but surely warrants discussion/disputes, clearly an evidence of karma farming!
RSS was removed from most sites (unfortunately) due to business, not technical reasons. Just like Web 2.0 era "mashup" friendly APIs, business started locking down their data despite it being an extremely user-hostile move.
I feel obligated to mention ledger[0] and hledger[1], those are plain text accounting software, well, as you can guess from their names, they allow you to do personal accounting in plain text.
Tangential to the actual point of the post, but the talk "Plain text? Really?" by Dylan Beattie[1] is one of my favorite talks. It does a great job capturing the problems with something that "seems" so simple. I think the author of the blog is well aware of these, though :)
Many things I knew, and a bunch I didn't. Thanks for that.
I wish he'd spent a little more time on the fundamental issue of CR vs CR/LF vs LF. That's been a "plain" text nightmare since before Unicode and many of the other complexities existed.
There was a challenge awhile back to describe a picture in 1000 words can't remember where I saw it. The winner just made a encoder/decoder to turn jpeg into words. The image was somewhat compressed but contained more visual data then you could accurately represent in the same way a raw description could. I thought it was so cool.
Graydon Hoare, the creator of the Rust programming language, wrote his seminal piece on text in 2014. In it, he said "text is the most powerful, useful, effective communication technology ever, period."[1] Text is durable.
[1] https://archive.ph/FhG5L (the original either got deleted or login-walled, here is the archived version)
[2] https://hn.algolia.com/?q=always+bet+on+text
[3] https://news.ycombinator.com/item?id=26164001
[4] https://news.ycombinator.com/item?id=8451271
[5] https://news.ycombinator.com/item?id=10284202
[6] https://news.ycombinator.com/item?id=12815829
I love plain text, but the calculus of being able to access the content in 50 years is much more interesting in the age of agents. They can infer the meaning of structured data without schemas, decompressed archives, find embedded files, etc. It's not perfect, but the durability of binary formats is better now than it has ever been.
It's not baked-in to the weights of any model, but agents can write their own tools to work with arbitrary binary formats and get many of the benefits of off-the-shelf unix utilities.
I love that the Godot game engine has a textual scene description. That's a stark contrast to Unreal Engine's binary format for blueprints (visual scripting).
This is the public dividend of a standard finally winning. The article gives credit to Unicode, but it is the fact that ASCII unambiguously won that gives plain text its portability and longevity. It looks like Unicode is on its way to winning in the same way, but it is not there yet. Most text files I write are still pure ASCII because that's the only way to avoid unexpected glitches [0].
[0] Windows newlines not withstanding.
Ahh, plain text, wherein "plain" does some heavy lifting. If plain text had something as simple as a 4 byte signature it could have been soooo much better.
As it is programs have to guess the following;
A) encoding. If ANSI which code page? If unicode which encoding? In the case of utf-16 big or little endian?
B) line endings? CR? LF? CRLF? I guess no-one uses LFCR... right?
C) number formats? 100.000 or 100,000? Date formats? is that mm-dd-yyyy or dd-mm-yyyy?
D) what human-language is it in?
E) CSV? Don't get me started...
Yes. Text is a lot easier to load, parse, make guesses about than say XLS. Yes all of the above things can be "guessed" to a greater or lesser extent.
Yes BOM exists to at least try solving the encoding question. Of course most files don't use them. Lots of tools don't support them. And they don't solve any of the other issues.
But sure, ASCII files in English with US date formats, and Windows line endings....no problems at all...
And file endings..
If you write files in ASCII you’re already writing in UTF-8.
That’s true, though if you’re writing UTF-8 you have to handle the corner cases when reading UTF-8, of which there are many.
Fortunately they’re easy to test for and most languages have standard libraries that make this painless.
ASCII has, in principal, infinitely many superset.
And in practice, still a rather large number of them, going back to the 1970's - https://en.wikipedia.org/wiki/Extended_ASCII
What, no mention of our old friend, ASCII-armored Base64? For shame! Everything can be text with Base64, including things that have absolutely no business being text! Best of all, in light of popular widespread abuse of every available resource, Base64 is comparatively efficient! Bring on the petabytes! Yay and I'm not being completely sarcastic
ASCII (7-Bit) is the only widely understood charset there is. Everything beyond this point depends on the loaded charset.
The fact that unicode maps the lower 7 bits to its own character set is a nice touch but none of the unicode sets are plain text.
Unicode are multibyte characters with variable byte length and endianess at play. If you read it wrong or guess the length wrong your results might be anything but useful.
Fun timing on this plaintext conversation. I built a little browser based plaintext playwriting app this weekend for a little weekend project. Found myself really hating 1. how clunk screenwriting software can be and 2. How unsharable the files are.
With plaintext you can hand the file off to anyone on any device--the caveat being absolutely no one wants to be handed a plaintext script. The software I used ten years ago to write plays is long since deprecated and those files are basically unopenable. Plaintext however remains.
John August has also talked about using plain text for movie/television screenplays, specifying a format based on markdown and hollywood standard: https://fountain.io/
Sharing aside, I imagine a plain text format would also be helpful for version control.
@dang @tomhow
My understanding is that HN has started incorporating AI tooling in the back-end to scan for LLM-generated submissions and source content, in an effort to discourage it and encourage human-made content (with human discussions, one would hope). Why do we keep having these kinds of articles every day? Entire "apps" are generated - see the lighthouse one - and submitted, and are plain-as-day LLM-generated nonsense.
Anyway, flagged.
Where did you get that understanding from? I might have missed something. They penalize LLM comments. Do they also penalize submissions? Any evidence of that?
dang mentioned it in passing some weeks back (going off of memory). Also please publish new NoH/Kongor build.
I think you’ve posted on an unrelated topic.
No, this is the right article.
Doesn’t read like AI.
Even has a simple an/a mistake that I doubt a chatbot is going to make.
they don't get @ mentions, you're better off emailing hn@ycombinator.com
I skimmed this whole article and didn't see any classic LLM tells about it. Why do you think it is LLM generated?
The recently released models have improved their writing styles. I don't know about this article, but if that's any indication of their writing habits, the author's github pages repo is full of posts generated by Claude.
The whole structure had a faint smell to it. It's hard to put into word, there was no obvious "it's not the quiet part. It's the load bearing bedrock", but the way the thing flowed made me strongly suspect LLM involvement.
perused the posts and incidentally all are touching high-octane topics, posted only recently but surely warrants discussion/disputes, clearly an evidence of karma farming!
The is real is real.
It is littered with little tells like that.
I think text is getting a big boost from the age of ai.
Probably because plain text can encode any form of distilled data. Technically our DNA can be plain text lol
Prediction: RSS is coming back stronger than ever
RSS was removed from most sites (unfortunately) due to business, not technical reasons. Just like Web 2.0 era "mashup" friendly APIs, business started locking down their data despite it being an extremely user-hostile move.
.plan is really simpler syndication.
I feel obligated to mention ledger[0] and hledger[1], those are plain text accounting software, well, as you can guess from their names, they allow you to do personal accounting in plain text.
0: https://ledger-cli.org
1: https://hledger.org
I prefer the simpler beancount: https://github.com/beancount/beancount
All of these use slightly incompatible formats, the irony of which will not escape the educated reader.
Tangential to the actual point of the post, but the talk "Plain text? Really?" by Dylan Beattie[1] is one of my favorite talks. It does a great job capturing the problems with something that "seems" so simple. I think the author of the blog is well aware of these, though :)
[1] https://www.youtube.com/watch?v=_mZBa3sqTrI
Many things I knew, and a bunch I didn't. Thanks for that.
I wish he'd spent a little more time on the fundamental issue of CR vs CR/LF vs LF. That's been a "plain" text nightmare since before Unicode and many of the other complexities existed.
that's why gno.land contracts render to markdown as the standard. you can browse the world through your terminal.
imagine the world wide web but markdown (and decentralized). gno.land is that.
"A picture is worth a thousand words"
There was a challenge awhile back to describe a picture in 1000 words can't remember where I saw it. The winner just made a encoder/decoder to turn jpeg into words. The image was somewhat compressed but contained more visual data then you could accurately represent in the same way a raw description could. I thought it was so cool.
Sometimes I'd rather have ten unambiguous words.