UTF-8 bit by bit (2001)

栏目: IT技术 · 发布时间: 5年前

内容简介:Richard Suchenwirth 2001-02-28 - From a delightful debugging chat at theTcl chatroom, I was brought to write down what I think on UTF-8 analysis (cf.Unicode and UTF-8, see also).I imagine a UTF-8 string as a railroad. It operates single-unitThe xxxxx bits

Richard Suchenwirth 2001-02-28 - From a delightful debugging chat at theTcl chatroom, I was brought to write down what I think on UTF-8 analysis (cf.Unicode and UTF-8, see also).

I imagine a UTF-8 string as a railroad. It operates single-unit railcars (one-byte ASCII characters, to be known from the fact that the highest bit is 0), and trains (sequences of two or more bytes that together form a character). Each train consists of exactly one locomotive (you see I'm European) and one or more trailers . The locomotive indicates the length of the train, including itself, in the highest bits that form a consecutive row of 1's, and one 0 bit. Examples:

0xxxxxxx : I'm a railcar, just a single unit
 110xxxxx : I'm leading a train of length 2

The xxxxx bits are used for other purposes (locomotives carry some freight too ;-)

Trailers indicate that they are trailers by the initial bit sequence 10. This way, they can't be mistaken for railcars or locomotives. E.g.

10yyyyyy: I'm a trailer

The freight of the train is in the x's and y's. In that concrete case, a C program reported to have received the bytes C3 and A4. Written as binary, that's

Now for clarity we delimit the indicators with parens:

(110)00011 (10)100100

and can just remove them:

00011 100100
 11100100 => E4, the iso8859-1 value for German ä ("a umlaut").

The generalized rule for the indicator of each byte is "those bits from highest(leftmost) down, up to and including the first zero bit".

Now going the other way. In orthodox UTF-8, a NUL byte (\x00) is represented by a NUL byte. Plain enough. But in Tcl we sometimes want NUL bytes inside "binary" strings (e.g. image data), without them terminating it as a real NUL byte does. To represent a NUL byte without any physical NUL bytes, we treat it like a character above ASCII, which must be a minimum two bytes long:

(110)00000 (10)000000 => C0 80

Whoops. Took us a while, but now we can read UTF-8, bit by bit.

andrewsh 2010-03-12 - Please note that 0xc0 0x80 sequence is illegal in the "Real" UTF-8: [], []


以上就是本文的全部内容,希望本文的内容对大家的学习或者工作能带来一定的帮助,也希望大家多多支持 码农网

查看所有标签

猜你喜欢:

本站部分资源来源于网络,本站转载出于传递更多信息之目的,版权归原作者或者来源机构所有,如转载稿涉及版权问题,请联系我们

算法帝国

算法帝国

克里斯托弗•斯坦纳 / 李筱莹 / 人民邮电出版社 / 2014-6 / 49.00

人类正在步入与机器共存的科幻世界?看《纽约时报》畅销书作者讲述算法和机器学习技术如何悄然接管人类社会,带我们走进一个算法统治的世界。 今天,算法涉足的领域已经远远超出了其创造者的预期。特别是进入信息时代以后,算法的应用涵盖金融、医疗、法律、体育、娱乐、外交、文化、国家安全等诸多方面,显现出源于人类而又超乎人类的强大威力。本书是《纽约时报》畅销书作者的又一力作,通过一个又一个引人入胜的故事,向......一起来看看 《算法帝国》 这本书的介绍吧!

在线进制转换器
在线进制转换器

各进制数互转换器

随机密码生成器
随机密码生成器

多种字符组合密码

HTML 编码/解码
HTML 编码/解码

HTML 编码/解码