如何用Selenium抓取网页表格?当前遭遇TimeoutException异常
MLB网页表格抓取超时问题解决方案
问题根源
原代码触发TimeoutException,大概率是这几个原因:
- ChromeDriver路径转义错误,导致驱动加载异常
- 等待条件不合理(表格无需点击,用
element_to_be_clickable没必要) - 页面懒加载,表格未在等待时间内渲染完成
修正后的完整代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException import pandas as pd # 修复路径转义问题,用原始字符串避免转义错误 driver = webdriver.Chrome(r'C:\Drivers\chromedriver.exe') waitWD = WebDriverWait(driver, 20) # 延长等待时间到20秒 link = "https://fantasyteamadvice.com/dfs/mlb/ownership" driver.get(link) try: # 改用等待元素可见,更适合表格抓取场景 table = waitWD.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "table[data-testid='ownershipTablemlb']"))) # 滚动到表格位置,触发懒加载(如果需要) driver.execute_script("arguments[0].scrollIntoView();", table) html = table.get_attribute("outerHTML") data = pd.read_html(html)[0] data.to_csv("inputs/dk_ownership.csv", index=False) # 去掉索引列更整洁 print("表格抓取成功,已保存为dk_ownership.csv") except TimeoutException: print("超时错误:表格元素未在规定时间内加载") finally: driver.quit()
关键修改说明
- 路径转义修复:Windows路径中的反斜杠会被Python识别为转义字符,用
r'路径'原始字符串可以避免这个问题 - 等待条件调整:把
element_to_be_clickable换成visibility_of_element_located,因为我们只需要表格可见,不需要点击它 - 延长等待时间:从10秒增加到20秒,给页面足够的渲染时间
- 滚动触发加载:通过JS脚本滚动到表格位置,解决部分网站的懒加载问题
- 异常捕获:增加
try-except块,便于排查错误,确保浏览器正常关闭
额外排查建议
如果还是报错,检查这几点:
- 确认ChromeDriver版本和本地Chrome浏览器版本一致,版本不兼容会导致驱动失效
- 打开目标页面手动查看,是否需要登录、验证,或者表格选择器是否已变更
- 可以尝试添加
driver.implicitly_wait(10),设置全局隐式等待,辅助页面元素加载
内容的提问来源于stack exchange,提问作者terrell.bradford
相关产品推荐
相关产品推荐

