You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup抓取网站首个表格失败,代码输出异常求助

问题:无法正确抓取目标表格数据

我尝试抓取网页https://web.archive.org/web/20230908091635/https://en.wikipedia.org/wiki/List_of_largest_banks中的第一个表格(10个市值最高的相关条目表格),但编写的Python代码输出结果不符合预期,以下是我的代码和错误输出:

原代码

# Code for ETL operations on Country-GDP data
# Importing the required libraries
import pandas as pd
import requests
from bs4 import BeautifulSoup


url="https://web.archive.org/web/20230908091635/https://en.wikipedia.org/wiki/List_of_largest_banks"

# Function to extract data using read_html
def extract(url):
    ''' This function aims to extract the required
    information from the website and save it to a data frame. The
    function returns the data frame for further processing. '''
    response = requests.get(url).text
       
    soup = BeautifulSoup(response, 'html.parser')
    df = pd.DataFrame(soup)
    tables = soup.find_all('tbody')
    rows = tables[0].find_all('tr')
    for row in rows:
        col = row.find_all('td')    
            
    return df

# Call the function and print the DataFrame
df = extract(url)
print(df)

错误输出

0
0                                               html
1                                                 

2  [
, [[], 
, [], 
, [window.RufflePlayer=win...
3  
     FILE ARCHIVED ON 09:16:35 Sep 08, 2023 ...
4                                                 

5  
playback timings (ms):
  captures_list: 0.9...

问题分析

原代码存在两个核心问题:

  1. 直接将BeautifulSoup对象转换为DataFrame,这完全不符合pandas的使用逻辑——BeautifulSoup对象是HTML解析后的树形结构,无法直接转换为表格数据。
  2. 遍历表格行时仅获取了列元素,但没有提取其中的文本内容,也没有将这些内容添加到DataFrame中,等于白遍历。

修正方案

这里提供两种简单有效的解决方法:

方法一:使用pandas的read_html直接读取表格(推荐)

pandas的read_html可以自动识别网页中的表格,直接返回DataFrame列表,操作最简单:

import pandas as pd

url = "https://web.archive.org/web/20230908091635/https://en.wikipedia.org/wiki/List_of_largest_banks"

# 读取网页中的所有表格,返回DataFrame列表
tables = pd.read_html(url)
# 第一个表格就是目标表格(10个市值最高的条目)
target_df = tables[0]
# 打印前10行验证
print(target_df.head(10))

方法二:使用BeautifulSoup手动解析后构建DataFrame

如果需要更精细的控制,可以手动解析HTML表格内容:

import pandas as pd
import requests
from bs4 import BeautifulSoup

url = "https://web.archive.org/web/20230908091635/https://en.wikipedia.org/wiki/List_of_largest_banks"

def extract(url):
    response = requests.get(url).text
    soup = BeautifulSoup(response, 'html.parser')
    
    # 找到第一个表格的tbody
    table = soup.find('table')
    tbody = table.find('tbody')
    rows = tbody.find_all('tr')
    
    # 初始化存储数据的列表
    data = []
    # 提取表头
    headers = [th.text.strip() for th in rows[0].find_all('th')]
    
    # 遍历数据行
    for row in rows[1:]:
        cols = row.find_all('td')
        # 提取每列的文本并去掉多余空格
        row_data = [col.text.strip() for col in cols]
        data.append(row_data)
    
    # 构建DataFrame
    df = pd.DataFrame(data, columns=headers)
    return df

df = extract(url)
# 打印目标的前10条数据
print(df.head(10))

内容的提问来源于stack exchange,提问作者MACAVELI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 15:24:51