You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium抓取网页表格?当前遭遇TimeoutException异常

MLB网页表格抓取超时问题解决方案

问题根源

原代码触发TimeoutException,大概率是这几个原因:

  • ChromeDriver路径转义错误,导致驱动加载异常
  • 等待条件不合理(表格无需点击,用element_to_be_clickable没必要)
  • 页面懒加载,表格未在等待时间内渲染完成

修正后的完整代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException
import pandas as pd

# 修复路径转义问题,用原始字符串避免转义错误
driver = webdriver.Chrome(r'C:\Drivers\chromedriver.exe')
waitWD = WebDriverWait(driver, 20)  # 延长等待时间到20秒
link = "https://fantasyteamadvice.com/dfs/mlb/ownership"  
driver.get(link)

try:
    # 改用等待元素可见,更适合表格抓取场景
    table = waitWD.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "table[data-testid='ownershipTablemlb']")))
    # 滚动到表格位置,触发懒加载(如果需要)
    driver.execute_script("arguments[0].scrollIntoView();", table)
    html = table.get_attribute("outerHTML")
    data = pd.read_html(html)[0]
    data.to_csv("inputs/dk_ownership.csv", index=False)  # 去掉索引列更整洁
    print("表格抓取成功,已保存为dk_ownership.csv")
except TimeoutException:
    print("超时错误:表格元素未在规定时间内加载")
finally:
    driver.quit()

关键修改说明

  • 路径转义修复:Windows路径中的反斜杠会被Python识别为转义字符,用r'路径'原始字符串可以避免这个问题
  • 等待条件调整:把element_to_be_clickable换成visibility_of_element_located,因为我们只需要表格可见,不需要点击它
  • 延长等待时间:从10秒增加到20秒,给页面足够的渲染时间
  • 滚动触发加载:通过JS脚本滚动到表格位置,解决部分网站的懒加载问题
  • 异常捕获:增加try-except块,便于排查错误,确保浏览器正常关闭

额外排查建议

如果还是报错,检查这几点:

  • 确认ChromeDriver版本和本地Chrome浏览器版本一致,版本不兼容会导致驱动失效
  • 打开目标页面手动查看,是否需要登录、验证,或者表格选择器是否已变更
  • 可以尝试添加driver.implicitly_wait(10),设置全局隐式等待,辅助页面元素加载

内容的提问来源于stack exchange,提问作者terrell.bradford

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 23:22:02