You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup和pandas提取维基百科表格?DataFrame为空排查

维基百科表格提取后DataFrame为空的问题分析与修复

问题原因

  • 表格匹配错误:目标页面的主表格class是wikitable sortable,原代码仅匹配wikitable,会抓取到页面内其他无关小表格,这些表格的表头与数据行结构和目标表格不符,导致后续无法正确提取数据。
  • 行处理逻辑疏漏:table.find_all("tr")会包含表头行(带<th>的行),这类行没有<td>元素,循环时cells为空,再加上len(cells) == len(headers)的判断,直接跳过了所有有效数据行。
  • 已弃用方法使用:df.append()在pandas 2.0及以上版本已被弃用,虽不是导致空DataFrame的直接原因,但会触发警告,不推荐使用。

修复方案

修正后的BeautifulSoup版本代码

from bs4 import BeautifulSoup
import requests
import pandas as pd

url = "https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue"
page = requests.get(url)
soup = BeautifulSoup(page.text, "html.parser")
# 匹配正确的表格class,目标表格带sortable标识
table = soup.find("table", {"class": "wikitable sortable"})

headers = [header.text.strip() for header in table.find_all("th")]

df = pd.DataFrame(columns=headers)

# 跳过表头行,从第二行开始遍历数据行
rows = table.find_all("tr")[1:]
for row in rows:
    cells = row.find_all("td")
    cells = [cell.text.strip() for cell in cells]
    if cells:
        # 用pd.concat替代已弃用的append方法
        df = pd.concat([df, pd.Series(cells, index=headers).to_frame().T], ignore_index=True)

print(df)

更简便的pandas原生方法

pandas自带了维基百科表格提取的快捷方式,一行代码就能完成:

import pandas as pd

url = "https://en.wikipedia.org/wiki/List_of_largest_companies_by_revenue"
# 读取页面内所有表格,第一个就是目标表格
df = pd.read_html(url)[0]
print(df)

内容的提问来源于stack exchange,提问作者Shadrach Bobby

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 17:23:19