You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用C# HTMLAgilityPack提取多个HTML表格表头?代码问题排查

问题原因

你的问题出在XPath选择器的使用上:table.SelectNodes("//th")里的//th是全局选择器,它会从整个HTML文档的根节点开始查找所有<th>元素,而不是仅在当前遍历的<table>节点范围内查找。所以每次循环时,都会把页面上所有表格的表头单元格全部遍历一遍,最终headerCount会累计到页面所有<th>的总数(也就是你看到的61)。

修复方案

把XPath选择器改成.//th,其中点号(.)表示从当前节点(即当前的<table>)开始,向下查找所有后代节点中的<th>元素:

namespace DataCollection
{
    internal class Program
    {
        static void Main(string[] args)
        {
            int headerCount;
            HtmlWeb web = new HtmlWeb();
            HtmlDocument doc = web.Load("https://en.wikipedia.org/wiki/List_of_actors_who_have_played_the_Doctor");
            //Extracting the tables from the HTML
            foreach (HtmlNode table in doc.DocumentNode.SelectNodes("//table"))
            {
                headerCount = 0;
                //Extracting the header cells from each table
                foreach (HtmlNode headerCol in table.SelectNodes(".//th"))
                {
                    headerCount++;
                    Console.WriteLine(headerCount);
                }
             
            };
           
            Console.ReadLine();
        }
    }
}

如果你的目标表格都有<thead>标签,也可以用更精确的选择器./thead//th,这样只会查找当前表格表头区域内的<th>,避免误选表格主体里可能存在的<th>元素。

内容的提问来源于stack exchange,提问作者Abdul

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 18:56:14