You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用C#的HtmlAgilityPack提取HTML页面td元素中的所有href链接

使用HtmlAgilityPack提取深层表格中td内的href链接

我正尝试用HtmlAgilityPack提取HTML页面中所有td标签内的href链接,这些表格嵌套在深层结构里,没法直接获取tr下的所有td。每个表格外层都有class为"table-group"的父div,或许可以从这里入手?现在最大的问题是目标结构上方有很多父元素,想跳过它们直接解析目标部分。


HTML结构示例

<table>
    <thead>
    </thead>
    <tbody>
        <tr>
            <td><a href="https://path-to-pdf1" target="_blank">Link 1</a></td>
            <td><a href="https://path-to-pdf1" target="_blank">1</a></td>
        </tr>
        <tr>
            <td><a href="https://path-to-pdf2" target="_blank">Link 2</a></td>
            <td><a href="https://path-to-pdf2" target="_blank">2</a></td>
        </tr>
        <tr>
            <td><a href="https://path-to-pdf3" target="_blank">Link 3</a></td>
            <td><a href="https://path-to-pdf3" target="_blank">3</a></td>
        </tr>
    </tbody>
</table>

<table>
    <thead>
    </thead>
    <tbody>
        <tr>
            <td><a href="https://path-to-pdf4" target="_blank">Link 4</a></td>
            <td><a href="https://path-to-pdf4" target="_blank">4</a></td>
        </tr>
        <tr>
            <td><a href="https://path-to-pdf5" target="_blank">Link 5</a></td>
            <td><a href="https://path-to-pdf5" target="_blank">5</a></td>
        </tr>
        <tr>
            <td><a href="https://path-to-pdf6" target="_blank">Link 6</a></td>
            <td><a href="https://path-to-pdf6" target="_blank">6</a></td>
        </tr>
    </tbody>
</table>

期望结果

提取到去重后的链接:

  • https://path-to-pdf1
  • https://path-to-pdf2
  • https://path-to-pdf3
  • https://path-to-pdf4
  • https://path-to-pdf5
  • https://path-to-pdf6

尝试的代码

var html = @"https://myurl.com";

HtmlWeb web = new HtmlWeb();

var htmlDoc = web.Load(html);

var nodes = htmlDoc.DocumentNode.SelectNodes("//table/tbody/tr/td/a[0]");

foreach (var item in nodes)
{
    Console.WriteLine(item.Attributes["href"].Value);
}

Console.ReadKey();

修正后的解决方案

问题分析

原代码存在两个核心问题:

  1. XPath索引从1开始,a[0]会匹配不到任何元素;
  2. 未利用外层table-group容器缩小范围,可能误匹配无关表格,且未处理重复链接。

修正代码

var url = @"https://myurl.com";

HtmlWeb web = new HtmlWeb();
var htmlDoc = web.Load(url);

// 定位到目标容器下的所有表格内的a标签
var nodes = htmlDoc.DocumentNode.SelectNodes("//div[@class='table-group']//table//td//a");

if (nodes != null)
{
    // HashSet自动去重
    var uniqueLinks = new HashSet<string>();
    foreach (var item in nodes)
    {
        var href = item.GetAttributeValue("href", string.Empty);
        if (!string.IsNullOrEmpty(href) && uniqueLinks.Add(href))
        {
            Console.WriteLine(href);
        }
    }
}

Console.ReadKey();

代码说明

  1. 精准定位://div[@class='table-group']//table//td//a 先锁定目标表格的外层容器,再递归匹配所有嵌套的td内的a标签,跳过上层无关结构;
  2. 空值保护:添加nodes != null判断,避免空引用异常;
  3. 安全取属性:用GetAttributeValue代替直接访问Attributes,避免属性不存在时出错;
  4. 自动去重:通过HashSet过滤重复链接,符合期望结果。

如果不需要依赖外层table-group容器,可改用通用XPath://table//td//a,但可能匹配到页面其他无关表格的链接。


内容的提问来源于stack exchange,提问作者Blake Rivell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 16:50:28