HtmlAgilityPack - How to set custom encoding when

2019-02-21 03:51发布

Is it possible to set custom encoding when loading pages with the method below?

HtmlWeb hwWeb = new HtmlWeb();
HtmlDocument hd = hwWeb.load("myurl");

I want to set encoding to "iso-8859-9".

I use C# 4.0 and WPF.

Edit: The question has been answered on MSDN.

3条回答
老娘就宠你
2楼-- · 2019-02-21 04:14

A decent answer is over here which handles auto-detecting the encoding as well as some other nifty features:

C# and HtmlAgilityPack encoding problem

查看更多
我想做一个坏孩纸
3楼-- · 2019-02-21 04:35
var document = new HtmlDocument();

using (var client = new WebClient())
{
    using (var stream = client.OpenRead(url))
    {
        var reader = new StreamReader(stream, Encoding.GetEncoding("iso-8859-9"));
        var html = reader.ReadToEnd();
        document.LoadHtml(html);
    }
}

This is a simple version of the solution answered here (for some reasons it got deleted)

查看更多
孤傲高冷的网名
4楼-- · 2019-02-21 04:38

I suppose you could try overriding the encoding in the HtmlWeb object.

Try this:

var web = new HtmlWeb
{
    AutoDetectEncoding = false,
    OverrideEncoding = myEncoding,
};
var doc = web.Load(myUrl);

Note: It appears that the OverrideEncoding property was added to HTML agility pack in revision 76610 so it is not available in the current release v1.4 (66017). The next best thing to do would be to read the page manually with the encodings overridden.

查看更多
登录 后发表回答