使用Selenium爬取JODI Data网站无法获取全部列名求助
问题描述
我正在使用Selenium爬取JODI Data网站的数据,目前仅能获取前几列列名,后续列需横向滚动后HTML详情及列名才会更新。已尝试多种横向滚动方法,但输出仍未改变,仅能获取前14个列名。
以下是我的代码:
chrome_service = webdriver.ChromeService(executable_path=r'C:\Users\user\Downloads\chromedriver-win64\chromedriver-win64\chromedriver.exe') driver = webdriver.Chrome(service=chrome_service, options=chrome_options) driver.get(r'http://www.jodidb.org/TableViewer/tableView.aspx?ReportId=93905') time.sleep(5) col_names = driver.find_elements(By.XPATH, '//th[@class="TVItemColHeader"]') for c in col_names: print(c.get_attribute('outerHTML')) col_lists = [] col_lists = [c.text for c in col_names]
当前仅能获取的列名outerHTML输出:
<th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a0">Jan2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a1">Feb2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a2">Mar2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a3">Apr2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a4">May2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a5">Jun2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a6">Jul2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a7">Aug2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a8">Sep2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a9">Oct2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a10">Nov2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a11">Dec2002</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a12">Jan2003</th> <th class="TVItemColHeader" scope="col" style="border-bottom-style: none" id="a13">Feb2003</th>
解决方案
这个网站的表格采用动态渲染机制,只有当列进入可视区域后才会生成对应的HTML元素。常规滚动方法可能没定位到正确的滚动容器,导致无法触发新列加载,可按以下步骤解决:
- 定位正确的滚动容器:通过开发者工具查看元素,该页面的表格滚动容器为
div#TVOuterContainer(带有overflow-x: auto样式,承载横向滚动逻辑)。 - 循环滚动加载所有列:用JavaScript控制容器滚动到最右侧,每次滚动后等待新列渲染,直到滚动位置不再变化(说明已到容器最右端)。
修改后的完整代码:
from selenium import webdriver from selenium.webdriver.common.by import By import time chrome_service = webdriver.ChromeService(executable_path=r'C:\Users\user\Downloads\chromedriver-win64\chromedriver-win64\chromedriver.exe') driver = webdriver.Chrome(service=chrome_service, options=chrome_options) driver.get(r'http://www.jodidb.org/TableViewer/tableView.aspx?ReportId=93905') time.sleep(5) # 定位表格的滚动容器 scroll_container = driver.find_element(By.ID, "TVOuterContainer") # 循环滚动直到无法再滚动 last_scroll_pos = 0 while True: # 执行JS将容器滚动到最右侧 driver.execute_script("arguments[0].scrollLeft = arguments[0].scrollWidth", scroll_container) time.sleep(1) # 等待新列渲染 current_scroll_pos = driver.execute_script("return arguments[0].scrollLeft", scroll_container) if current_scroll_pos == last_scroll_pos: break last_scroll_pos = current_scroll_pos # 获取全部列名 col_names = driver.find_elements(By.XPATH, '//th[@class="TVItemColHeader"]') col_lists = [c.text for c in col_names] # 输出结果 print(f"共获取到 {len(col_lists)} 列列名:") print(col_lists) driver.quit()
- 关键注意事项:如果
TVOuterContainer不是正确的滚动容器,可通过开发者工具检查元素的overflow-x属性,找到实际承载横向滚动的父元素;等待时间可根据网络情况调整,确保新列完全渲染。
内容的提问来源于stack exchange,提问作者Stackie
相关产品推荐
相关产品推荐

