如何用BeautifulSoup和pandas提取维基百科表格?DataFrame为空排查
维基百科表格提取后DataFrame为空的问题分析与修复
问题原因
- 表格匹配错误:目标页面的主表格class是
wikitable sortable,原代码仅匹配wikitable,会抓取到页面内其他无关小表格,这些表格的表头与数据行结构和目标表格不符,导致后续无法正确提取数据。 - 行处理逻辑疏漏:
table.find_all("tr")会包含表头行(带<th>的行),这类行没有<td>元素,循环时cells为空,再加上len(cells) == len(headers)的判断,直接跳过了所有有效数据行。 - 已弃用方法使用:
df.append()在pandas 2.0及以上版本已被弃用,虽不是导致空DataFrame的直接原因,但会触发警告,不推荐使用。
修复方案
修正后的BeautifulSoup版本代码
from bs4 import BeautifulSoup import requests import pandas as pd url = "https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue" page = requests.get(url) soup = BeautifulSoup(page.text, "html.parser") # 匹配正确的表格class,目标表格带sortable标识 table = soup.find("table", {"class": "wikitable sortable"}) headers = [header.text.strip() for header in table.find_all("th")] df = pd.DataFrame(columns=headers) # 跳过表头行,从第二行开始遍历数据行 rows = table.find_all("tr")[1:] for row in rows: cells = row.find_all("td") cells = [cell.text.strip() for cell in cells] if cells: # 用pd.concat替代已弃用的append方法 df = pd.concat([df, pd.Series(cells, index=headers).to_frame().T], ignore_index=True) print(df)
更简便的pandas原生方法
pandas自带了维基百科表格提取的快捷方式,一行代码就能完成:
import pandas as pd url = "https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue" # 读取页面内所有表格,第一个就是目标表格 df = pd.read_html(url)[0] print(df)
内容的提问来源于stack exchange,提问作者Shadrach Bobby
相关产品推荐
相关产品推荐

