You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium限制获取表格前9列?解决Pandas数据长度不匹配报错

解决表格转Pandas DataFrame的两种方案

方案一:只提取前9列(最直接,符合你的需求)

既然你只需要前9列数据,直接在获取表头和行数据时做切片即可,完全避开后面的问题列。修改代码如下:

for x in result:
    driver.find_element(By.XPATH, '//*[@id="sidemenu"]/table/tbody/tr[1]/td/form/div[2]/input[1]').send_keys(x)
    driver.implicitly_wait(2)
    driver.find_element(By.XPATH, '//*[@id="navsrch"]').click()
    driver.implicitly_wait(2)

    headers = []
    columns = dict()

    table_id = driver.find_element(By.ID, 'invoice')
    all_rows = table_id.find_elements(By.TAG_NAME, "tr")

    row = all_rows[0]
    # 只取前9个表头
    all_items = row.find_elements(By.TAG_NAME, "th")[:9]
    for item in all_items:
        name = item.text
        columns[name] = []
        headers.append(name)

    print(headers)

    for row in all_rows[1:]:
        # 只取每行的前9个单元格
        all_items = row.find_elements(By.TAG_NAME, "td")[:9]
        for name, item in zip(headers, all_items):
            value = item.text
            columns[name].append(value)

    print(columns)
    df = pd.DataFrame(columns)
    print(df)

driver.close()

核心改动就是在获取th和td后加上[:9]切片,只保留前9列数据,空表头、重复表头的问题都不会影响到你。

方案二:修复表头问题(如果需要完整15列数据时用)

如果之后需要完整数据,可以通过处理空表头和重复表头来解决报错:

  • 空表头:自动生成类似Column_10、Column_11这样的命名
  • 重复表头:给重复的表头添加后缀(比如Name_1、Name_2)

修改代码如下:

for x in result:
    driver.find_element(By.XPATH, '//*[@id="sidemenu"]/table/tbody/tr[1]/td/form/div[2]/input[1]').send_keys(x)
    driver.implicitly_wait(2)
    driver.find_element(By.XPATH, '//*[@id="navsrch"]').click()
    driver.implicitly_wait(2)

    headers = []
    columns = dict()
    # 用来记录表头出现的次数,处理重复
    header_count = {}

    table_id = driver.find_element(By.ID, 'invoice')
    all_rows = table_id.find_elements(By.TAG_NAME, "tr")

    row = all_rows[0]
    all_items = row.find_elements(By.TAG_NAME, "th")
    for idx, item in enumerate(all_items):
        name = item.text.strip()
        # 处理空表头
        if not name:
            name = f"Column_{idx+1}"
        # 处理重复表头
        if name in header_count:
            header_count[name] += 1
            name = f"{name}_{header_count[name]}"
        else:
            header_count[name] = 0
        columns[name] = []
        headers.append(name)

    print(headers)

    for row in all_rows[1:]:
        all_items = row.find_elements(By.TAG_NAME, "td")
        # 确保列数匹配,防止某行数据缺失
        if len(all_items) != len(headers):
            # 补全缺失的单元格为空字符串
            all_items += [""] * (len(headers) - len(all_items))
        for name, item in zip(headers, all_items):
            value = item.text
            columns[name].append(value)

    print(columns)
    df = pd.DataFrame(columns)
    print(df)

driver.close()

这样处理后,所有表头都会是唯一且非空的,同时还处理了行数据列数不匹配的情况,避免ValueError报错。

内容的提问来源于stack exchange,提问作者Xylis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 18:48:23