Comment by josephg
14 hours ago
I disagree. UTF-8 is the right default. Pure ASCII strings are rarely needed in modern software. Unless you know that you need ASCII, your strings should support unicode.
14 hours ago
I disagree. UTF-8 is the right default. Pure ASCII strings are rarely needed in modern software. Unless you know that you need ASCII, your strings should support unicode.
Internet Messages have been ASCII since Internet Messages.
True, but humans have been speaking non English languages since humans.
The first networked computers predate unicode by several decades. Back then, the word length could be 8 but not neccesarily. All sorts of different encodings existed and a given common standard wasn't yet agreed upon
It's not like internet standards don't know there are other languages, it's just that they documented how things were done at the time. Some legacy has remained ever since.
Most internet protocols these days have shifted to UTF-8.
Perhaps so, and that's a good development imho. However, the most notorious of all, HTTP, has not afaik. Different encodings have been suggested including UTF-8. I'm not entirely sure why those proposals weren't implemented, but the dot-com bubble at the time might've had something to do with it.