You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取HTML表格无法存入Pandas DataFrame?求错误排查

Selenium爬取HDI表格转Pandas DataFrame报错修复

错误原因

报错ValueError: Shape of passed values is (9, 1), indices imply (9, 9)的核心问题:

  • 数据收集逻辑错误:每次遍历表格行时都重置row_data,最终仅保留最后一行的数据,且row_data是一维列表,无法匹配表头的9列结构
  • CSS选择器错误:使用tr, td会同时选中行元素本身和单元格,导致数据混入无效内容

修复后的完整代码

from selenium import webdriver 
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.service import Service
import time
import pandas as pd

def HDI():
    url = 'https://worldpopulationreview.com/country-rankings/hdi-by-country'

    service = Service(executable_path="C:/driver/new/chromedriver_win32/chromedriver.exe")
    driver = webdriver.Chrome(service=service)
    driver.get(url)
    time.sleep(5)

    btn = driver.find_element(By.CLASS_NAME, '_3p_1XEZR')
    btn.click()
    time.sleep(5)

    temp_height=0

    while True:
        # 滚动页面加载全部数据
        driver.execute_script("window.scrollBy(0,500)")
        time.sleep(5)
        check_height = driver.execute_script("return document.documentElement.scrollTop || window.pageYOffset || document.body.scrollTop;")
        if check_height==temp_height:
           break
        temp_height=check_height
    time.sleep(3)

    # 提取表头
    row_headers = []
    tableheads = driver.find_elements(By.CLASS_NAME, 'datatable-th')
    for value in tableheads:
        thead_values = value.find_element(By.CLASS_NAME, 'has-tooltip-bottom').text.strip()
        row_headers.append(thead_values)

    # 提取表格数据
    all_rows_data = []  # 存储所有行数据的二维列表
    tablebodies = driver.find_elements(By.TAG_NAME, 'tr')
    for row in tablebodies:
        # 仅选择当前行的td元素,排除tr本身
        tabledata = row.find_elements(By.CSS_SELECTOR, 'td')
        if tabledata:  # 跳过没有单元格的空行
            row_data = []
            for data in tabledata:
                row_data.append(data.text.strip())
            all_rows_data.append(row_data)

    # 创建DataFrame
    df = pd.DataFrame(all_rows_data, columns=row_headers)
    print(df.head())  # 打印前5行验证
    # 可选:保存为CSV
    # df.to_csv('hdi_data.csv', index=False)
    
    driver.quit()  # 关闭浏览器

HDI()

关键修改说明

  • 数据存储结构调整:新增all_rows_data作为二维列表,每次遍历行时将当前行的row_data添加进去,确保所有行数据都被收集
  • 修正CSS选择器:将tr, td改为td,只提取单元格数据,避免混入行元素的无效内容
  • 空行过滤:添加if tabledata:判断,跳过没有单元格的空行,避免生成空数据行
  • 关闭浏览器:新增driver.quit(),避免浏览器进程残留

内容的提问来源于stack exchange,提问作者Miracle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 10:50:15