使用BeautifulSoup抓取网站首个表格失败,代码输出异常求助
问题:无法正确抓取目标表格数据
我尝试抓取网页https://web.archive.org/web/20230908091635/https://en.wikipedia.org/wiki/List_of_largest_banks中的第一个表格(10个市值最高的相关条目表格),但编写的Python代码输出结果不符合预期,以下是我的代码和错误输出:
原代码
# Code for ETL operations on Country-GDP data # Importing the required libraries import pandas as pd import requests from bs4 import BeautifulSoup url="https://web.archive.org/web/20230908091635/https://en.wikipedia.org/wiki/List_of_largest_banks" # Function to extract data using read_html def extract(url): ''' This function aims to extract the required information from the website and save it to a data frame. The function returns the data frame for further processing. ''' response = requests.get(url).text soup = BeautifulSoup(response, 'html.parser') df = pd.DataFrame(soup) tables = soup.find_all('tbody') rows = tables[0].find_all('tr') for row in rows: col = row.find_all('td') return df # Call the function and print the DataFrame df = extract(url) print(df)
错误输出
0 0 html 1 2 [ , [[], , [], , [window.RufflePlayer=win... 3 FILE ARCHIVED ON 09:16:35 Sep 08, 2023 ... 4 5 playback timings (ms): captures_list: 0.9...
问题分析
原代码存在两个核心问题:
- 直接将
BeautifulSoup对象转换为DataFrame,这完全不符合pandas的使用逻辑——BeautifulSoup对象是HTML解析后的树形结构,无法直接转换为表格数据。 - 遍历表格行时仅获取了列元素,但没有提取其中的文本内容,也没有将这些内容添加到
DataFrame中,等于白遍历。
修正方案
这里提供两种简单有效的解决方法:
方法一:使用pandas的read_html直接读取表格(推荐)
pandas的read_html可以自动识别网页中的表格,直接返回DataFrame列表,操作最简单:
import pandas as pd url = "https://web.archive.org/web/20230908091635/https://en.wikipedia.org/wiki/List_of_largest_banks" # 读取网页中的所有表格,返回DataFrame列表 tables = pd.read_html(url) # 第一个表格就是目标表格(10个市值最高的条目) target_df = tables[0] # 打印前10行验证 print(target_df.head(10))
方法二:使用BeautifulSoup手动解析后构建DataFrame
如果需要更精细的控制,可以手动解析HTML表格内容:
import pandas as pd import requests from bs4 import BeautifulSoup url = "https://web.archive.org/web/20230908091635/https://en.wikipedia.org/wiki/List_of_largest_banks" def extract(url): response = requests.get(url).text soup = BeautifulSoup(response, 'html.parser') # 找到第一个表格的tbody table = soup.find('table') tbody = table.find('tbody') rows = tbody.find_all('tr') # 初始化存储数据的列表 data = [] # 提取表头 headers = [th.text.strip() for th in rows[0].find_all('th')] # 遍历数据行 for row in rows[1:]: cols = row.find_all('td') # 提取每列的文本并去掉多余空格 row_data = [col.text.strip() for col in cols] data.append(row_data) # 构建DataFrame df = pd.DataFrame(data, columns=headers) return df df = extract(url) # 打印目标的前10条数据 print(df.head(10))
内容的提问来源于stack exchange,提问作者MACAVELI
相关产品推荐
相关产品推荐

