You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取JODI Data网站无法获取全部列名求助

问题描述

我正在使用Selenium爬取JODI Data网站的数据,目前仅能获取前几列列名,后续列需横向滚动后HTML详情及列名才会更新。已尝试多种横向滚动方法,但输出仍未改变,仅能获取前14个列名。

以下是我的代码:

chrome_service = webdriver.ChromeService(executable_path=r'C:\Users\user\Downloads\chromedriver-win64\chromedriver-win64\chromedriver.exe')
driver = webdriver.Chrome(service=chrome_service, options=chrome_options)
driver.get(r'http://www.jodidb.org/TableViewer/tableView.aspx?ReportId=93905')


time.sleep(5)

col_names = driver.find_elements(By.XPATH, '//th[@class="TVItemColHeader"]')


for c in col_names:
    print(c.get_attribute('outerHTML'))

col_lists = []

col_lists = [c.text for c in col_names]

当前仅能获取的列名outerHTML输出:

<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a0">Jan2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a1">Feb2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a2">Mar2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a3">Apr2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a4">May2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a5">Jun2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a6">Jul2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a7">Aug2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a8">Sep2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a9">Oct2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a10">Nov2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a11">Dec2002</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a12">Jan2003</th>
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a13">Feb2003</th>
解决方案

这个网站的表格采用动态渲染机制,只有当列进入可视区域后才会生成对应的HTML元素。常规滚动方法可能没定位到正确的滚动容器,导致无法触发新列加载,可按以下步骤解决:

  • 定位正确的滚动容器:通过开发者工具查看元素,该页面的表格滚动容器为div#TVOuterContainer(带有overflow-x: auto样式,承载横向滚动逻辑)。
  • 循环滚动加载所有列:用JavaScript控制容器滚动到最右侧,每次滚动后等待新列渲染,直到滚动位置不再变化(说明已到容器最右端)。

修改后的完整代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
import time

chrome_service = webdriver.ChromeService(executable_path=r'C:\Users\user\Downloads\chromedriver-win64\chromedriver-win64\chromedriver.exe')
driver = webdriver.Chrome(service=chrome_service, options=chrome_options)
driver.get(r'http://www.jodidb.org/TableViewer/tableView.aspx?ReportId=93905')

time.sleep(5)

# 定位表格的滚动容器
scroll_container = driver.find_element(By.ID, "TVOuterContainer")

# 循环滚动直到无法再滚动
last_scroll_pos = 0
while True:
    # 执行JS将容器滚动到最右侧
    driver.execute_script("arguments[0].scrollLeft = arguments[0].scrollWidth", scroll_container)
    time.sleep(1)  # 等待新列渲染
    current_scroll_pos = driver.execute_script("return arguments[0].scrollLeft", scroll_container)
    if current_scroll_pos == last_scroll_pos:
        break
    last_scroll_pos = current_scroll_pos

# 获取全部列名
col_names = driver.find_elements(By.XPATH, '//th[@class="TVItemColHeader"]')
col_lists = [c.text for c in col_names]

# 输出结果
print(f"共获取到 {len(col_lists)} 列列名:")
print(col_lists)

driver.quit()
  • 关键注意事项:如果TVOuterContainer不是正确的滚动容器,可通过开发者工具检查元素的overflow-x属性,找到实际承载横向滚动的父元素;等待时间可根据网络情况调整,确保新列完全渲染。

内容的提问来源于stack exchange,提问作者Stackie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 11:39:51