如何用Selenium和C#获取网站动态加载的隐藏产品表格数据?
解决Selenium动态加载表格提取问题
爬取网址http://www.ifm.com/de/en/category/200_010_010_010的产品表格数据时,因表格是JavaScript/AJAX动态加载,原代码无法定位元素,所有定位表达式返回null。原代码如下:
var options = new ChromeOptions() { BinaryLocation = "C:\Program Files (x86)\Google\Chrome\Application\chrome.exe", }; options.AddArguments(new List<string>() { "headless", "disable-gpu" }); string response = ""; options.AddArgument("no-sandbox"); using (var browser = new ChromeDriver(options)) { browser.Navigate().GoToUrl(url); WebDriverWait wait = new WebDriverWait(browser, TimeSpan.FromSeconds(20)); /// /// *Both below expresions return null* //IWebElement rows_count = browser.FindElement(By.XPath("ifm-selector__matching-products")); //IWebElement next_button = browser.FindElement(By.XPath("ifm-pagination__cta normalize hover- link-2")); response= browser.PageSource; } HtmlAgilityPack.HtmlDocument htmlDoc = new HtmlAgilityPack.HtmlDocument(); htmlDoc.LoadHtml(response); var rows_count = htmlDoc.DocumentNode.SelectSingleNode("//div[@class='ifm- selector__results']//div[@class='ifm-selector__matching-products']//span");
问题根源
- XPath定位语法错误:原代码中XPath直接写类名,未通过
@class属性定位元素,语法完全无效。 - 未等待动态内容渲染:页面跳转后立即获取
PageSource,此时动态表格还未加载完成,DOM中不存在目标元素。 - 冗余解析流程:Selenium已持有渲染完成的浏览器DOM,无需导出
PageSource再用HtmlAgilityPack解析,多此一举且易丢失动态内容。
解决方法
方案1:直接用Selenium获取表格及行数
修正定位语法,等待目标元素加载完成后直接提取数据:
var options = new ChromeOptions() { BinaryLocation = @"C:\Program Files (x86)\Google\Chrome\Application\chrome.exe", }; options.AddArguments(new List<string>() { "headless", "disable-gpu", "no-sandbox" }); using (var browser = new ChromeDriver(options)) { browser.Navigate().GoToUrl("http://www.ifm.com/de/en/category/200_010_010_010"); WebDriverWait wait = new WebDriverWait(browser, TimeSpan.FromSeconds(20)); // 等待匹配产品数元素加载完成 IWebElement matchingProducts = wait.Until(driver => driver.FindElement(By.CssSelector("div.ifm-selector__matching-products span"))); // 获取产品行数文本 string rowCountText = matchingProducts.Text; Console.WriteLine($"产品行数信息:{rowCountText}"); // 等待表格加载完成,获取表格元素 IWebElement productTable = wait.Until(driver => driver.FindElement(By.CssSelector("table.ifm-table"))); // 获取表格所有行 IList<IWebElement> tableRows = productTable.FindElements(By.TagName("tr")); Console.WriteLine($"表格实际行数:{tableRows.Count}"); }
方案2:等待后用HtmlAgilityPack解析
若必须使用HtmlAgilityPack,需确保等待动态内容加载完成后再获取PageSource:
var options = new ChromeOptions() { BinaryLocation = @"C:\Program Files (x86)\Google\Chrome\Application\chrome.exe", }; options.AddArguments(new List<string>() { "headless", "disable-gpu", "no-sandbox" }); string response = ""; using (var browser = new ChromeDriver(options)) { browser.Navigate().GoToUrl("http://www.ifm.com/de/en/category/200_010_010_010"); WebDriverWait wait = new WebDriverWait(browser, TimeSpan.FromSeconds(20)); // 等待关键元素加载,确保动态表格渲染完成 wait.Until(driver => driver.FindElement(By.CssSelector("div.ifm-selector__matching-products"))); // 此时获取PageSource才包含动态内容 response = browser.PageSource; } HtmlAgilityPack.HtmlDocument htmlDoc = new HtmlAgilityPack.HtmlDocument(); htmlDoc.LoadHtml(response); // 用contains处理类名中的空格问题 var rows_count = htmlDoc.DocumentNode.SelectSingleNode("//div[contains(@class,'ifm-selector__matching-products')]//span"); if (rows_count != null) { Console.WriteLine($"产品行数:{rows_count.InnerText}"); } // 获取表格行数 var tableRows = htmlDoc.DocumentNode.SelectNodes("//table[contains(@class,'ifm-table')]//tr"); if (tableRows != null) { Console.WriteLine($"表格行数:{tableRows.Count}"); }
关键注意事项
- 优先使用CSS选择器定位元素,比XPath更简洁不易出错;若用XPath,类名含空格时需用
contains处理(避免DOM中类名顺序/合并导致的定位失败)。 - 必须用
WebDriverWait的Until方法等待元素可见/可交互,不要用固定延时,适配不同网络和渲染速度。 - 无头模式下若出现加载失败,可添加
--window-size=1920,1080参数模拟正常窗口尺寸,规避反爬限制。
内容的提问来源于stack exchange,提问作者S.Parseh
相关产品推荐
相关产品推荐

