PHP Parsing Problem - and Â

When I try to parse some html that has   sprinkled through it and then echo it, the   "turns into" this character: Â. Also, html_entity_decode() and str_replace() doesn't change it.

Why is this happening? How can I remove the Â's?

标签： php html parsing character-encoding

3条回答

forever°为你锁心

2楼-- · 2019-01-11 07:17

The non-breaking space exist in UTF-8 of two bytes: 0xC2 and 0xA0.

When those bytes are represented in ISO-8859-1 (a single-byte encoding) instead of UTF-8 (a multi-byte encoding) then those bytes becomes respectively the characters Â and another non-breaking space .

Apparently you're parsing the HTML using UTF-8 and echoing the results using ISO-8859-1. To fix this problem, you need to either parse HTML using ISO-8859-1 or echo the results using UTF-8. I'd recommend to use UTF-8 all the way. Go through the PHP UTF-8 cheatsheet to align it all out.

0人赞添加讨论(0) 举报

我想做一个坏孩纸

3楼-- · 2019-01-11 07:29

preg_replace() can also do the trick:

preg_replace("/&#?[a-z0-9]{2,8};/i","", $var);

0人赞添加讨论(0) 举报

可以哭但决不认输i

4楼-- · 2019-01-11 07:33

html_entity_decode("&nbsp;") == '\xa0'

I think by design, I don't understand why str_replace does not work for you, try this snippet:

$nbsp = html_entity_decode("&nbsp;");
$s = html_entity_decode("[&nbsp;]");
$s = str_replace($nbsp, " ", $s);
echo $s;

perhaps \xa0 it's not a valid unicode string, so using the result of the html_entity_decode() may be more appropriate for text replacement instead of \xa0.

BalusC explanation looks plausible you may trying to insert utf-8 \xc2\xa0 in the the then trying to display it as latin instead of utf8, if you want to use unicode stuff you should keep utf-8 encoding everywhere, from the charset of the server to the db, since you will have the same problem when using e.g. à

0人赞添加讨论(0) 举报

PHP Parsing Problem - and Â

采纳回答

编辑标签

举报内容

检举类型

检举原因

检举说明(必填)

打开微信“扫一扫”，打开网页后点击屏幕右上角分享按钮

付费偷看金额在0.1-10元之间