使用C#的HtmlAgilityPack提取HTML页面td元素中的所有href链接
使用HtmlAgilityPack提取深层表格中td内的href链接
我正尝试用HtmlAgilityPack提取HTML页面中所有td标签内的href链接,这些表格嵌套在深层结构里,没法直接获取tr下的所有td。每个表格外层都有class为"table-group"的父div,或许可以从这里入手?现在最大的问题是目标结构上方有很多父元素,想跳过它们直接解析目标部分。
HTML结构示例
<table> <thead> </thead> <tbody> <tr> <td><a href="https://path-to-pdf1" target="_blank">Link 1</a></td> <td><a href="https://path-to-pdf1" target="_blank">1</a></td> </tr> <tr> <td><a href="https://path-to-pdf2" target="_blank">Link 2</a></td> <td><a href="https://path-to-pdf2" target="_blank">2</a></td> </tr> <tr> <td><a href="https://path-to-pdf3" target="_blank">Link 3</a></td> <td><a href="https://path-to-pdf3" target="_blank">3</a></td> </tr> </tbody> </table> <table> <thead> </thead> <tbody> <tr> <td><a href="https://path-to-pdf4" target="_blank">Link 4</a></td> <td><a href="https://path-to-pdf4" target="_blank">4</a></td> </tr> <tr> <td><a href="https://path-to-pdf5" target="_blank">Link 5</a></td> <td><a href="https://path-to-pdf5" target="_blank">5</a></td> </tr> <tr> <td><a href="https://path-to-pdf6" target="_blank">Link 6</a></td> <td><a href="https://path-to-pdf6" target="_blank">6</a></td> </tr> </tbody> </table>
期望结果
提取到去重后的链接:
- https://path-to-pdf1
- https://path-to-pdf2
- https://path-to-pdf3
- https://path-to-pdf4
- https://path-to-pdf5
- https://path-to-pdf6
尝试的代码
var html = @"https://myurl.com"; HtmlWeb web = new HtmlWeb(); var htmlDoc = web.Load(html); var nodes = htmlDoc.DocumentNode.SelectNodes("//table/tbody/tr/td/a[0]"); foreach (var item in nodes) { Console.WriteLine(item.Attributes["href"].Value); } Console.ReadKey();
修正后的解决方案
问题分析
原代码存在两个核心问题:
- XPath索引从1开始,
a[0]会匹配不到任何元素; - 未利用外层
table-group容器缩小范围,可能误匹配无关表格,且未处理重复链接。
修正代码
var url = @"https://myurl.com"; HtmlWeb web = new HtmlWeb(); var htmlDoc = web.Load(url); // 定位到目标容器下的所有表格内的a标签 var nodes = htmlDoc.DocumentNode.SelectNodes("//div[@class='table-group']//table//td//a"); if (nodes != null) { // HashSet自动去重 var uniqueLinks = new HashSet<string>(); foreach (var item in nodes) { var href = item.GetAttributeValue("href", string.Empty); if (!string.IsNullOrEmpty(href) && uniqueLinks.Add(href)) { Console.WriteLine(href); } } } Console.ReadKey();
代码说明
- 精准定位:
//div[@class='table-group']//table//td//a先锁定目标表格的外层容器,再递归匹配所有嵌套的td内的a标签,跳过上层无关结构; - 空值保护:添加
nodes != null判断,避免空引用异常; - 安全取属性:用
GetAttributeValue代替直接访问Attributes,避免属性不存在时出错; - 自动去重:通过
HashSet过滤重复链接,符合期望结果。
如果不需要依赖外层table-group容器,可改用通用XPath://table//td//a,但可能匹配到页面其他无关表格的链接。
内容的提问来源于stack exchange,提问作者Blake Rivell
相关产品推荐
相关产品推荐

