如何用Selenium与Headless Chrome抓取网页表格?定位失败求助
问题描述
尝试编写Python脚本从Digikey产品页面导出Product Attributes表格为Excel/CSV文件,但脚本报错找不到目标元素。代码在其他网站可正常运行,但在Digikey、Mouser这类元器件网站上失效,怀疑被网站反爬机制拦截。
目标表格截图:
原代码
import time import requests from selenium import webdriver from selenium.webdriver.chrome.options import Options import pandas as pd def get_specifications_table(url): options = Options() options.add_argument('--headless') # 无头模式运行浏览器(无可视窗口) driver = webdriver.Chrome(options=options) driver.get(url) time.sleep(5) # 延迟等待页面加载(可按需调整时间) try: # 查找指定类名的元素并提取表格 class_name = "MuiTable-root css-u6unfi" table_element = driver.find_element("css selector", f".{class_name}") table_html = table_element.get_attribute('outerHTML') df = pd.read_html(table_html)[0] return df except Exception as e: print("Error:", e) finally: driver.quit() return None def export_to_excel(df, output_file): writer = pd.ExcelWriter(output_file, engine='xlsxwriter') df.to_excel(writer, index=False) writer.save() writer.close() if __name__ == '__main__': url = "https://www.digikey.com/en/products/detail/texas-instruments/uln2003aidre4/1912622" output_excel_file = "Specifications_Table_Digikey.xlsx" print("Fetching the webpage and extracting the table...") specifications_df = get_specifications_table(url) if specifications_df is not None: print("Exporting the table to Excel...") export_to_excel(specifications_df, output_excel_file) print(f"Table 'Specifications' exported to '{output_excel_file}' successfully.") else: print("Table extraction or export failed.")
报错信息
Fetching the webpage and extracting the table... Error: Message: no such element: Unable to locate element: {"method":"css selector","selector":".MuiTable-root.css-u6unfi"} (Session info: headless chrome=115.0.5790.110); For documentation on this error, please visit: https://www.selenium.dev/documentation/webdriver/troubleshooting/errors#no-such-element-exception Stacktrace: Backtrace: GetHandleVerifier [0x004BA813+48355] (No symbol) [0x0044C4B1] (No symbol) [0x00355358] (No symbol) [0x003809A5] (No symbol) [0x00380B3B] (No symbol) [0x003AE232] (No symbol) [0x0039A784] (No symbol) [0x003AC922] (No symbol) [0x0039A536] (No symbol) [0x003782DC] (No symbol) [0x003793DD] GetHandleVerifier [0x0071AABD+2539405] GetHandleVerifier [0x0075A78F+2800735] GetHandleVerifier [0x0075456C+2775612] GetHandleVerifier [0x005451E0+616112] (No symbol) [0x00455F8C] (No symbol) [0x00452328] (No symbol) [0x0045240B] (No symbol) [0x00444FF7] BaseThreadInitThunk [0x772500C9+25] RtlGetAppContainerNamedObjectPath [0x77BC7B4E+286] RtlGetAppContainerNamedObjectPath [0x77BC7B1E+238] Table extraction or export failed.
解决方案
1. 绕过反爬检测
Digikey会识别无头Chrome的自动化特征,需要添加参数模拟正常浏览器:
options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36') options.add_argument('--disable-blink-features=AutomationControlled') options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False)
2. 替换固定等待为显式等待
time.sleep(5)的固定等待不可靠,改用Selenium的显式等待确保元素加载完成:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 替换原有的time.sleep和元素查找逻辑 wait = WebDriverWait(driver, 15) # 用更稳定的属性定位(比如表格的aria-label) table_element = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'table[aria-label="Product Attributes"]')))
3. 修正CSS选择器错误
原代码中依赖的css-u6unfi是框架动态生成的类名,会随网站更新变化,必须改用固定属性或父容器定位,避免后续失效。
4. 完整修复后的代码
import time from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd def get_specifications_table(url): options = Options() options.add_argument('--headless=new') # 新版无头模式,更接近正常浏览器行为 options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36') options.add_argument('--disable-blink-features=AutomationControlled') options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option('useAutomationExtension', False) driver = webdriver.Chrome(options=options) # 隐藏自动化标识 driver.execute_script("Object.defineProperty(navigator, 'webdriver', {get: () => undefined})") try: driver.get(url) # 显式等待表格加载完成 wait = WebDriverWait(driver, 15) table_element = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, 'table[aria-label="Product Attributes"]'))) table_html = table_element.get_attribute('outerHTML') df = pd.read_html(table_html)[0] return df except Exception as e: print("错误:", e) finally: driver.quit() return None def export_to_excel(df, output_file): writer = pd.ExcelWriter(output_file, engine='xlsxwriter') df.to_excel(writer, index=False) writer.close() if __name__ == '__main__': url = "https://www.digikey.com/en/products/detail/texas-instruments/uln2003aidre4/1912622" output_excel_file = "Specifications_Table_Digikey.xlsx" print("获取网页并提取表格...") specifications_df = get_specifications_table(url) if specifications_df is not None: print("导出表格到Excel...") export_to_excel(specifications_df, output_excel_file) print(f"表格已成功导出到 '{output_excel_file}'") else: print("表格提取或导出失败。")
5. 额外建议
- Digikey提供官方API,直接调用API获取数据比网页爬取更稳定,还能避免网页结构变化导致的失效;
- 批量处理产品数据时,优先使用官方API;
- 永远不要依赖框架动态生成的类名(如
css-u6unfi)定位元素,这类名称会随时改变。
内容的提问来源于stack exchange,提问作者Arman Zamani
相关产品推荐
相关产品推荐

