Detecting character encoding in HTML

2020-02-11 01:24发布

I download an HTML page. The HTTP content-type header specifies one character encoding, and the page has a meta tag that specifies another. What's the correct way to handle that?

I guess 'correct' isn't the right word, since nobody follows the damn standards anyway... so what's the way that will cause me the least problems?

标签： html http character-encoding

1条回答

爷的心禁止访问

2楼-- · 2020-02-11 02:08

Do the same as webbrowsers do: use the response header. When HTML is served over HTTP, the meta tag is ignored when the response header is present. Only when the HTML is read from local disk file system, the meta tag is been used. This is also explicitly specified by w3 HTML spec.

To sum up, conforming user agents must observe the following priorities when determining a document's character encoding (from highest priority to lowest):

An HTTP "charset" parameter in a "Content-Type" field.

A META declaration with "http-equiv" set to "Content-Type" and a value set for "charset".

The charset attribute set on an element that designates an external resource.

Any existing decent HTML parser in whatever language you use should already take this into account. According your question history you're familiar with Java, I'd then suggest to grab Jsoup for this.

0人赞添加讨论(0) 举报

Detecting character encoding in HTML

采纳回答

编辑标签

举报内容

检举类型

检举原因

检举说明(必填)

打开微信“扫一扫”，打开网页后点击屏幕右上角分享按钮

付费偷看金额在0.1-10元之间