You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取遇Mismatched columns错误:无法设置行的解决方法

爬取维基百科企业列表时cannot set rows in Mismatched columns错误排查与修复

我在爬取维基百科「全球营收最高企业列表」页面的表格时,遇到了cannot set rows in Mismatched columns错误,代码如下:

from bs4 import BeautifulSoup
import requests #importing beautifulsoup and requests

url = 'https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue'

page = requests.get(url)

soup = BeautifulSoup(page.text, 'html') #storing beautiful soup in soup variable

print(soup) # to see what it displays

soup.find('table') # finding every tag labelled table

soup.find('table', class_ = 'wikitable sortable') # trying to specify tables

table = soup.find_all('table')[0] # this is the one I want

print(table) # to see what it displays

table.find_all('th') # find all the th (column headings) tags in the table.

world_titles = table.find_all('th') # so i can just type world_titles instead of table.find_all('th') all the time

world_table_titles = [title.text.strip() for title in world_titles] # removing /n and making data clean

print(world_table_titles) # seeing what it displays

import pandas as pd # importing pandas

df = pd.DataFrame(columns = world_table_titles) # making a dataframe

df

column_data = table.find_all('tr') # finding rows within my table

for row in column_data[2:]: # [2:] because the first two just displayed []
    row_data = row.find_all('td')
    individual_row_data = [data.text.strip() for data in row_data] #clean version of row_data
    
    length = len(df)  
    df.loc[length] = individual_row_data
    print(individual_row_data)

错误截图如下:
错误截图

我已在YouTube及网络寻找解决方案但未获帮助,请问问题原因是什么?该如何修复?


问题原因

  1. 列数不匹配:核心问题是DataFrame定义的列数(从所有<th>提取的表头)和部分数据行的实际列数不一致。维基百科的表格存在跨行表头(比如表头区域有多行<th>),导致你提取的表头列数多于实际数据行的列数。
  2. 行遍历逻辑错误:原代码从column_data[2:]开始遍历,但部分行可能因为合并单元格(rowspan/colspan)导致<td>数量不足,强行赋值到DataFrame就会触发列数不匹配错误。

修复方案

修改后的代码

from bs4 import BeautifulSoup
import requests
import pandas as pd

url = 'https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue'
page = requests.get(url)
soup = BeautifulSoup(page.text, 'html')

# 精准定位目标表格,避免取错表格
table = soup.find('table', class_='wikitable sortable')

# 只提取第一行的有效表头,避免跨行表头干扰
header_row = table.find('tr')
world_table_titles = [title.text.strip() for title in header_row.find_all('th')]

# 创建匹配列数的DataFrame
df = pd.DataFrame(columns=world_table_titles)

# 从第二行开始遍历数据行(跳过表头行)
for row in table.find_all('tr')[1:]:
    row_data = row.find_all('td')
    # 只处理列数和表头一致的行,跳过合并单元格的异常行
    if len(row_data) == len(world_table_titles):
        individual_row_data = [data.text.strip() for data in row_data]
        df.loc[len(df)] = individual_row_data

# 查看结果
print(df.head())

关键优化点

  • 精准定位表头:只提取表格第一行的<th>作为表头,确保列数和数据行匹配。
  • 过滤异常行:添加列数判断,跳过因合并单元格导致列数不足的行,避免赋值报错。
  • 可靠表格定位:直接用class_='wikitable sortable'定位目标表格,比find_all('table')[0]更稳定,避免页面结构变化导致取错表格。

内容的提问来源于stack exchange,提问作者Rango00

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 09:19:55