如何在C#中利用HtmlDocument和XPath根据Key-content获取Value-content?
问题描述
我在C#中使用HtmlDocument的LoadHtml方法加载了一个HTML页面,尝试遍历HTML结构。目标HTML片段如下:
<tr class="xxxx"> <td class="yyyy"> <span>Key-content</span> <sup class ="zzzz">abcd</sup> </td> <td class="wwww">Value-content</td> </tr>
我希望通过DocumentNode.SelectSingleNode方法,根据已知的Key-content字符串,获取对应的Value-content元素内容,但无法写出正确的XPath表达式。已尝试的代码如下:
Task<HttpResponseMessage> response = client.GetAsync(address, HttpCompletionOption.ResponseContentRead); using (HttpResponseMessage message = response.Result) { using (HttpContent content = message.Content) { string xmlData = content.ReadAsStringAsync().Result; HtmlDocument htmlDoc = new HtmlDocument(); htmlDoc.LoadHtml(xmlData); string xpath = "//*['Key-content']/../../td[1]"; // ==> //*['Key-content']: 定位包含'Key-content'的<span>节点 // ==> /../..: 向上回溯两次到<tr>节点 // ==> /td[1]: 定位目标节点 HtmlNode node = htmlDoc.DocumentNode.SelectSingleNode(xpath); string text = node.InnerText; } }
解决方案
原XPath的问题
- 文本匹配写法错误:
//*['Key-content']无法正确匹配包含指定文本的节点,需要用text()函数筛选,正确的文本匹配写法是//span[text()='Key-content'](因为文本明确在<span>里,直接指定标签更精准)。 - 目标节点索引错误:
<tr>下的第一个<td>是包含Key-content的节点,第二个才是存储Value-content的节点,所以应该用td[2]而非td[1]。
正确的XPath表达式
可以用两种写法实现:
- 基于父节点回溯:
//span[text()='Key-content']/../../td[2]
- 用
ancestor轴定位父级<tr>(可读性更强):
//span[text()='Key-content']/ancestor::tr/td[2]
修正后的代码
Task<HttpResponseMessage> response = client.GetAsync(address, HttpCompletionOption.ResponseContentRead); using (HttpResponseMessage message = response.Result) { using (HttpContent content = message.Content) { string xmlData = content.ReadAsStringAsync().Result; HtmlDocument htmlDoc = new HtmlDocument(); htmlDoc.LoadHtml(xmlData); // 使用更清晰的ancestor轴写法 string xpath = "//span[text()='Key-content']/ancestor::tr/td[2]"; HtmlNode node = htmlDoc.DocumentNode.SelectSingleNode(xpath); // 增加空节点判断,避免空引用异常 string text = node?.InnerText ?? string.Empty; } }
内容的提问来源于stack exchange,提问作者Douar Gwenn
相关产品推荐
相关产品推荐

