使用Selenium获取需点击按钮显示的表格数据遇阻
问题原因分析
- 工具方法混用错误:你用Selenium获取了
table元素,但后续调用的find('tbody')、find_all('tr')是BeautifulSoup的API,Selenium的WebElement对象不支持这些方法,两种爬虫工具的语法被混淆了。 - 缺少动态加载等待:点击"Show"按钮后表格是异步渲染的,没有等待表格完全加载就去提取数据,大概率拿不到有效内容。
- 元素定位语法错误:
driver.find_element('#state-snapshot-data-table')写法不规范,Selenium的find_element必须指定定位方式,这里应该搭配By.CSS_SELECTOR来使用这个选择器。 - 浏览器过早关闭:
finally块里直接执行driver.quit(),导致后续处理表格数据时,浏览器已经关闭,所有WebElement对象都失效了。
解决思路及修正代码
核心修正点
- 统一使用Selenium的元素操作方法,或者将Selenium获取的页面源码传给BeautifulSoup解析(二选一即可)。
- 点击按钮后,显式等待表格的核心元素(比如tbody)加载完成,确保数据渲染完毕。
- 修正元素定位的语法,遵循Selenium的API规范。
- 调整浏览器关闭时机,等数据提取完成后再执行关闭操作。
方案一:纯Selenium实现
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd def getData(state, year): url = f"https://www.countyhealthrankings.org/explore-health-rankings/county-health-rankings-model/health-outcomes/length-of-life/infant-mortality?year={year}&state={state}&tab=1" driver = webdriver.Chrome() driver.get(url) data = [] try: # 等待"Show"按钮可点击并点击 WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH, '//span[contains(text(), "Show")]'))).click() # 等待表格tbody加载完成 tbody = WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.CSS_SELECTOR, '#state-snapshot-data-table tbody'))) # 用Selenium方法遍历行和单元格 rows = tbody.find_elements(By.TAG_NAME, 'tr') for row in rows: cells = row.find_elements(By.TAG_NAME, 'td') if len(cells) >= 4: # 避免因单元格缺失报错 name = cells[0].text.strip() numerator = cells[1].text.strip() raw_value = cells[2].text.strip() ci_range = cells[3].text.strip() data.append([name, numerator, raw_value, ci_range]) except Exception as e: print(f"错误信息: {str(e)}") return pd.DataFrame() finally: driver.quit() return pd.DataFrame(data, columns=['name', 'numerator', 'raw_value', 'ci_range']) print( getData('01', '2023') )
方案二:Selenium+BeautifulSoup混合实现
如果更习惯用BeautifulSoup解析HTML结构,可以在表格加载完成后,提取页面源码交给BeautifulSoup处理:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd from bs4 import BeautifulSoup def getData(state, year): url = f"https://www.countyhealthrankings.org/explore-health-rankings/county-health-rankings-model/health-outcomes/length-of-life/infant-mortality?year={year}&state={state}&tab=1" driver = webdriver.Chrome() driver.get(url) data = [] try: WebDriverWait(driver, 10).until(EC.element_to_be_clickable((By.XPATH, '//span[contains(text(), "Show")]'))).click() # 等待表格整体加载完成 WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.ID, 'state-snapshot-data-table'))) # 获取页面源码,用BeautifulSoup解析 soup = BeautifulSoup(driver.page_source, 'html.parser') table = soup.find('table', id='state-snapshot-data-table') if table: tbody = table.find('tbody') rows = tbody.find_all('tr') for row in rows: cells = row.find_all('td') if len(cells) >= 4: name = cells[0].text.strip() numerator = cells[1].text.strip() raw_value = cells[2].text.strip() ci_range = cells[3].text.strip() data.append([name, numerator, raw_value, ci_range]) except Exception as e: print(f"错误信息: {str(e)}") return pd.DataFrame() finally: driver.quit() return pd.DataFrame(data, columns=['name', 'numerator', 'raw_value', 'ci_range']) print( getData('01', '2023') )
内容的提问来源于stack exchange,提问作者eth
相关产品推荐
相关产品推荐

