在python 2.7中打印UTF-8字符

Lin Ma 发表于 Dev

Lin Ma

这是我打开，阅读和输出的方式。该文件是用于Unicode字符的UTF-8编码文件。我想打印前10个UTF-8字符，但是下面代码段的输出显示了10个无法识别的怪异字符。想知道是否有人对如何正确打印有任何想法？谢谢。

   with open(name, 'r') as content_file:
        content = content_file.read()
        for i in range(10):
            print content[i]

10个怪异角色中的每个角色都看起来像这样，

�

问候，林

2号环

将Unicode代码点（字符）编码为UTF-8时，某些代码点将转换为单个字节，但是许多代码点会变成一个以上的字节。标准7位ASCII范围内的字符将被编码为单个字节，但是更多的外来字符通常将需要更多的字节进行编码。

因此，您将获得那些奇怪的字符，因为您将这些多字节UTF-8序列分解为单个字节。有时这些字节将对应于普通的可打印字符，但通常它们不会，因此您可以打印出来。

这是一个使用©，®和™字符的简短演示，这两个字符在UTF-8中分别被编码为2、2和3个字节。我的终端设置为使用UTF-8。

utfbytes = "\xc2\xa9 \xc2\xae \xe2\x84\xa2"
print utfbytes, len(utfbytes)
for b in utfbytes:
    print b, repr(b)

uni = utfbytes.decode('utf-8')
print uni, len(uni)

输出

Stack Overflow的联合创始人Joel Spolsky在Unicode上写了一篇很好的文章：每个软件开发人员绝对绝对要完全了解Unicode和字符集（绝对没有借口！）

您还应该查看Python文档中的Unicode HOWTO文章，以及Ned Batchelder的实用Unicode文章（又称为“ Unipain”）。

这是从UTF-8编码的字节字符串中提取单个字符的简短示例。正如我在评论中提到的那样，要正确执行此操作，您需要知道每个字符被编码为多少个字节。

utfbytes = "\xc2\xa9 \xc2\xae \xe2\x84\xa2"
widths = (2, 1, 2, 1, 3)
start = 0
for w in widths:
    print "%d %d [%s]" % (start, w, utfbytes[start:start+w])
    start += w

输出

0 2 [©]
2 1 [ ]
3 2 [®]
5 1 [ ]
6 3 [™]

FWIW，这是该代码的Python 3版本：

utfbytes = b"\xc2\xa9 \xc2\xae \xe2\x84\xa2"
widths = (2, 1, 2, 1, 3)
start = 0
for w in widths:
    s = utfbytes[start:start+w]
    print("%d %d [%s]" % (start, w, s.decode()))
    start += w

如果我们不知道UTF-8字符串中字符的字节宽度，那么我们需要做更多的工作。每个UTF-8序列都会在第一个字节中编码该序列的宽度，如Wikipedia文章中有关UTF-8所述。

以下Python 2演示演示了如何提取该宽度信息。它产生与前两个片段相同的输出。

# UTF-8 code widths
#width starting byte
#1 0xxxxxxx
#2 110xxxxx
#3 1110xxxx
#4 11110xxx
#C 10xxxxxx

def get_width(b):
    if b <= '\x7f':
        return 1
    elif '\x80' <= b <= '\xbf':
        #Continuation byte
        raise ValueError('Bad alignment: %r is a continuation byte' % b)
    elif '\xc0' <= b <= '\xdf':
        return 2
    elif '\xe0' <= b <= '\xef':
        return 3
    elif '\xf0' <= b <= '\xf7':
        return 4
    else:
        raise ValueError('%r is not a single byte' % b)


utfbytes = b"\xc2\xa9 \xc2\xae \xe2\x84\xa2"
start = 0
while start < len(utfbytes):
    b = utfbytes[start]
    w = get_width(b)
    s = utfbytes[start:start+w]
    print "%d %d [%s]" % (start, w, s)
    start += w

一般来说，它应该不会有必要做这样的事情：只使用所提供的解码方法。

出于好奇，这里是的Python 3版本get_width，以及一个手动解码UTF-8字节串的函数。

def get_width(b):
    if b <= 0x7f:
        return 1
    elif 0x80 <= b <= 0xbf:
        #Continuation byte
        raise ValueError('Bad alignment: %r is a continuation byte' % b)
    elif 0xc0 <= b <= 0xdf:
        return 2
    elif 0xe0 <= b <= 0xef:
        return 3
    elif 0xf0 <= b <= 0xf7:
        return 4
    else:
        raise ValueError('%r is not a single byte' % b)

def decode_utf8(utfbytes):
    start = 0
    uni = []
    while start < len(utfbytes):
        b = utfbytes[start]
        w = get_width(b)
        if w == 1:
            n = b
        else:
            n = b & (0x7f >> w)
            for b in utfbytes[start+1:start+w]:
                if not 0x80 <= b <= 0xbf:
                    raise ValueError('Not a continuation byte: %r' % b)
                n <<= 6
                n |= b & 0x3f
        uni.append(chr(n))
        start += w
    return ''.join(uni)


utfbytes = b'\xc2\xa9 \xc2\xae \xe2\x84\xa2'
print(utfbytes.decode('utf8'))
print(decode_utf8(utfbytes))

输出

©®™
©®™

本文收集自互联网，转载请注明来源。

如有侵权，请联系[email protected] 删除。

编辑于2021-03-3

我来说两句

0条评论

登录后参与评论

上一篇：如何使用pyspark和regex在字符串的RDD中查找所有以my_str开头的单词？

来自分类Dev

Related 相关文章

文章

在python 2.7中打印UTF-8字符

在python 2.7中打印UTF-8字符

在Python 3和Python 2中处理CSV中的非UTF8字符

无法转换UTF-8字符-Python

Python反转UTF-8字符串

用Python计算UTF8字符

Python：如何使用UTF-8字符进行打印？

如何从MySql在ZF2中显示utf8字符

如何在python中构建utf8字符串

utf-8字符串从python到AWS中的Java android

Swift 2 Json utf8字符串字符错误

如何删除python字符串的最后utf8字符

从Python Unicode字符串获取UTF-8字符代码

Python：将utf-8字符串转换为字节字符串

Python：将utf-8字符串转换为字节字符串

从Python Unicode字符串获取UTF-8字符代码

在Python 3中，如何从字符串中删除所有非UTF8字符？

在python 3中将转义的utf-8字符串转换为utf

带有UTF-8字符的HTML2canvas图像捕获问题

在Jinja2模板中使用utf-8字符

2个不同的UTF-8字符串上的MySQL UNIQUE错误？

ZXing 2D条码解码：UTF-8字符未正确解码

UTF-8无法在我的python代码中编码UTF-8字符。它们显示为原义UTF-8

如何在Python中用前面的数字分割utf-8字符串？

Python 3.3 C-API和UTF-8字符串

Python将UTF8字符串插入SQLite

清单元素上的python 3.4 UTF 8字符串

Python：如何从sqlite数据库查询utf-8字符串

在Python中将utf-8字符串拆分为字节

我必须使用哪种python编码类型来读取非utf-8字符？

Python 3.5.2-help（）函数未正确显示å，ä，ö（UTF-8字符）