网页爬取遇Mismatched columns错误:无法设置行的解决方法
爬取维基百科企业列表时
cannot set rows in Mismatched columns错误排查与修复 我在爬取维基百科「全球营收最高企业列表」页面的表格时,遇到了cannot set rows in Mismatched columns错误,代码如下:
from bs4 import BeautifulSoup import requests #importing beautifulsoup and requests url = 'https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue' page = requests.get(url) soup = BeautifulSoup(page.text, 'html') #storing beautiful soup in soup variable print(soup) # to see what it displays soup.find('table') # finding every tag labelled table soup.find('table', class_ = 'wikitable sortable') # trying to specify tables table = soup.find_all('table')[0] # this is the one I want print(table) # to see what it displays table.find_all('th') # find all the th (column headings) tags in the table. world_titles = table.find_all('th') # so i can just type world_titles instead of table.find_all('th') all the time world_table_titles = [title.text.strip() for title in world_titles] # removing /n and making data clean print(world_table_titles) # seeing what it displays import pandas as pd # importing pandas df = pd.DataFrame(columns = world_table_titles) # making a dataframe df column_data = table.find_all('tr') # finding rows within my table for row in column_data[2:]: # [2:] because the first two just displayed [] row_data = row.find_all('td') individual_row_data = [data.text.strip() for data in row_data] #clean version of row_data length = len(df) df.loc[length] = individual_row_data print(individual_row_data)
错误截图如下:
我已在YouTube及网络寻找解决方案但未获帮助,请问问题原因是什么?该如何修复?
问题原因
- 列数不匹配:核心问题是DataFrame定义的列数(从所有
<th>提取的表头)和部分数据行的实际列数不一致。维基百科的表格存在跨行表头(比如表头区域有多行<th>),导致你提取的表头列数多于实际数据行的列数。 - 行遍历逻辑错误:原代码从
column_data[2:]开始遍历,但部分行可能因为合并单元格(rowspan/colspan)导致<td>数量不足,强行赋值到DataFrame就会触发列数不匹配错误。
修复方案
修改后的代码
from bs4 import BeautifulSoup import requests import pandas as pd url = 'https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue' page = requests.get(url) soup = BeautifulSoup(page.text, 'html') # 精准定位目标表格,避免取错表格 table = soup.find('table', class_='wikitable sortable') # 只提取第一行的有效表头,避免跨行表头干扰 header_row = table.find('tr') world_table_titles = [title.text.strip() for title in header_row.find_all('th')] # 创建匹配列数的DataFrame df = pd.DataFrame(columns=world_table_titles) # 从第二行开始遍历数据行(跳过表头行) for row in table.find_all('tr')[1:]: row_data = row.find_all('td') # 只处理列数和表头一致的行,跳过合并单元格的异常行 if len(row_data) == len(world_table_titles): individual_row_data = [data.text.strip() for data in row_data] df.loc[len(df)] = individual_row_data # 查看结果 print(df.head())
关键优化点
- 精准定位表头:只提取表格第一行的
<th>作为表头,确保列数和数据行匹配。 - 过滤异常行:添加列数判断,跳过因合并单元格导致列数不足的行,避免赋值报错。
- 可靠表格定位:直接用
class_='wikitable sortable'定位目标表格,比find_all('table')[0]更稳定,避免页面结构变化导致取错表格。
内容的提问来源于stack exchange,提问作者Rango00
相关产品推荐
相关产品推荐

