urllib2 opener providing wrong charset

2020-02-06 08:04发布

When I open the url and read it, I can't recognize it. But when I check the content header it says it is encoded as utf-8. So I tried to convert it to unicode and it complained UnicodeDecodeError: 'ascii' codec can't decode byte 0x8b in position 1: ordinal not in range(128) using unicode().

.encode("utf-8") produces UnicodeDecodeError: 'ascii' codec can't decode byte 0x8b in position 1: ordinal not in range(128)

.decode("utf-8") produced UnicodeDecodeError: 'utf8' codec can't decode byte 0x8b in position 1: invalid start byte.

I have tried everything I can come up with(I'm not that good at encodings)

I would be happy if I could get this to work. Thanks.

标签： python utf-8 character-encoding urllib2

2条回答

男人必须洒脱

2楼-- · 2020-02-06 08:17

The header is probably wrong. Check out chardet.

EDIT: Thinking more about it -- my money is on the contents being gzipped. I believe some of Python's various URL-opening modules/classes/etc will ungzip, while others won't.

0人赞添加讨论(0) 举报

你好瞎i

3楼-- · 2020-02-06 08:20

This is a common mistake. The server sends gzipped stream.

You should unpack it first:

response = opener.open(self.__url, data)
if response.info().get('Content-Encoding') == 'gzip':
    buf = StringIO.StringIO( response.read())
    gzip_f = gzip.GzipFile(fileobj=buf)
    content = gzip_f.read()
else:
    content = response.read()

0人赞添加讨论(0) 举报

urllib2 opener providing wrong charset

采纳回答

编辑标签

举报内容

检举类型

检举原因

检举说明(必填)

打开微信“扫一扫”，打开网页后点击屏幕右上角分享按钮

付费偷看金额在0.1-10元之间