如何用System.Xml.XmlDocument解析HTML?转换及解析报错求助
Let's break down what's going on here and how to fix these issues so you can get that HTML converted to valid XHTML for your C# XML parsing:
1. Invalid Character Warnings (Codes 128 and 153)
The warnings about replacing invalid characters (128 and 153) stem from encoding mismatches or non-XML-compliant characters:
- Code 128 is a control character that’s not allowed in standard XML.
- Code 153 is the trademark symbol (™), a Windows-1252 character that might not be properly encoded in your input HTML.
Fixes:
- Specify the input encoding when running tidy. If your HTML uses Windows-1252, use this command:
tidy -asxhtml -m -encoding windows-1252 index.html - Preserve raw characters instead of replacing them (if you need to keep symbols like ™):
tidy -asxhtml -m -raw index.html - Pre-process in C#: Before running tidy, clean the HTML to remove or replace invalid XML characters. For example, filter out disallowed control characters:
string cleanedHtml = new string(originalHtml.Where(c => XmlConvert.IsXmlChar(c)).ToArray());
2. Unrecognized Content Errors (Lines 121, 125)
The "无法识别!" (unrecognized) errors mean tidy hit malformed HTML syntax it can’t fix automatically—this is almost always broken tags or invalid structure.
Fixes:
- Inspect the problematic lines: Open
index.htmland check lines 121 and 125 for common issues like:- Unclosed tags (e.g.,
<div>without a matching</div>) - Missing quotes around attributes (e.g.,
<div class=foo>instead of<div class="foo">) - Custom elements tidy doesn’t recognize (use
-new-blocklevel-tagsor-new-inline-tagsto whitelist them if needed)
- Unclosed tags (e.g.,
- Manually correct the syntax: Fix the broken HTML in those lines, then re-run tidy—these errors should vanish.
3. Verify the Final XHTML for C# Parsing
After resolving tidy’s warnings/errors, confirm the output XHTML is well-formed (a must for C# XML parsing):
- Use C#'s
XmlReaderto test parsing directly:using (XmlReader reader = XmlReader.Create("index.html")) { while (reader.Read()) { // Iterate through to catch any remaining parsing errors } } - If errors persist, double-check for unclosed tags or leftover invalid characters tidy might have missed.
Remember: HTML is built to be forgiving, but XML (and XHTML) is strict. Tidy is a powerful tool, but it can’t fix every edge case—manual checks for broken syntax are often necessary.
内容的提问来源于stack exchange,提问作者Don't Know

