You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium进行Python Web Scraping时仅爬取表格首行数据的问题求助

解决Selenium爬取表格仅拿到第一行数据的问题

嘿,我看了你的代码,马上发现了几个导致你只拿到第一行数据的问题,咱们一步步来修正:

核心问题分析

  1. 硬编码了行索引tr[1]:你所有的XPath都固定指向了表格的第一行(tr[1]),这自然只会获取第一行的数据,完全没遍历其他行。
  2. 没有利用循环的表格对象:你在循环tables的时候,却直接用了table[2]这种固定定位,等于白循环了所有表格,一直只操作第二个表格。
  3. 全局查找元素:用driver.find_elements是从整个页面找元素,应该基于当前表格/行来做相对查找,避免干扰其他元素。

修正后的代码

from selenium.common.exceptions import NoSuchElementException
import time

driver = visit_main_page()
contents = driver.find_elements_by_xpath('//*[@id="mw-content-text"]/div[1]')
# 用相对路径查找当前div下的表格,避免全局匹配
tables = contents[0].find_elements_by_xpath('./table')
data = {"Date": [], "Time": [], "Place": [], "Latitude": [], "Longitude": [], "Fatalities": [], "Magnitude": []}

for table in tables:
    try:
        # 获取当前表格下的所有行,tbody是浏览器自动补的,有的页面可能没有,可以改成./tr
        rows = table.find_elements_by_xpath('./tbody/tr')
        # 跳过表头行(假设第一行是表头,根据实际页面调整索引)
        for row in rows[1:]:
            # 基于当前行做相对查找,定位每一列
            date_col = row.find_element_by_xpath('./td[1]')
            time_col = row.find_element_by_xpath('./td[2]')
            place_col = row.find_element_by_xpath('./td[3]')
            lat_col = row.find_element_by_xpath('./td[4]')
            long_col = row.find_element_by_xpath('./td[5]')
            fat_col = row.find_element_by_xpath('./td[6]')
            magn_col = row.find_element_by_xpath('./td[7]')
            
            # 清理文本并追加到数据字典
            data['Date'].append(date_col.text.strip())
            data['Time'].append(time_col.text.strip())
            data['Place'].append(place_col.text.strip())
            data['Latitude'].append(lat_col.text.strip())
            data['Longitude'].append(long_col.text.strip())
            data['Fatalities'].append(fat_col.text.strip())
            data['Magnitude'].append(magn_col.text.strip())
    except NoSuchElementException:
        print('当前行缺少部分字段,跳过该行')
        continue
    time.sleep(1)

关键修改点说明

  • 相对路径查找:用table.find_elements和row.find_element代替全局的driver.find_elements,确保只在当前表格/行范围内找元素。
  • 遍历所有数据行:去掉了固定的tr[1],改成遍历rows[1:](跳过表头,如果你不需要跳过表头就去掉[1:])。
  • 直接处理单行列数据:每个行对应一组数据,直接提取每列文本后追加,不需要额外嵌套循环。
  • 文本清理:用.strip()去掉文本前后的空格、换行符,让数据更整洁。

额外优化建议

  • 别用time.sleep,改用WebDriverWait等待元素加载,更高效也更稳定:
    from selenium.webdriver.support.ui import WebDriverWait
    from selenium.webdriver.support import expected_conditions as EC
    from selenium.webdriver.common.by import By
    
    # 等待目标区域加载完成
    WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.XPATH, '//*[@id="mw-content-text"]/div[1]/table')))
    
  • 如果只需要处理第二个表格,不用循环所有tables,直接取tables[1](Python是0索引)即可,节省不必要的循环。
  • 检查页面实际HTML结构,有些表格可能没有<tbody>标签,这时候把XPath里的./tbody/tr改成./tr就行。

内容的提问来源于stack exchange,提问作者Marianna Karavangeli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 22:27:37