如何用C# HTMLAgilityPack提取多个HTML表格表头?代码问题排查
问题原因
你的问题出在XPath选择器的使用上:table.SelectNodes("//th")里的//th是全局选择器,它会从整个HTML文档的根节点开始查找所有<th>元素,而不是仅在当前遍历的<table>节点范围内查找。所以每次循环时,都会把页面上所有表格的表头单元格全部遍历一遍,最终headerCount会累计到页面所有<th>的总数(也就是你看到的61)。
修复方案
把XPath选择器改成.//th,其中点号(.)表示从当前节点(即当前的<table>)开始,向下查找所有后代节点中的<th>元素:
namespace DataCollection { internal class Program { static void Main(string[] args) { int headerCount; HtmlWeb web = new HtmlWeb(); HtmlDocument doc = web.Load("https://en.wikipedia.org/wiki/List_of_actors_who_have_played_the_Doctor"); //Extracting the tables from the HTML foreach (HtmlNode table in doc.DocumentNode.SelectNodes("//table")) { headerCount = 0; //Extracting the header cells from each table foreach (HtmlNode headerCol in table.SelectNodes(".//th")) { headerCount++; Console.WriteLine(headerCount); } }; Console.ReadLine(); } } }
如果你的目标表格都有<thead>标签,也可以用更精确的选择器./thead//th,这样只会查找当前表格表头区域内的<th>,避免误选表格主体里可能存在的<th>元素。
内容的提问来源于stack exchange,提问作者Abdul
相关产品推荐
相关产品推荐

