如何抓取同名多表格?解决BeautifulSoup多表格抓取报错问题
解决多表格抓取及Excel导出问题
错误原因
soup.find_all('table')返回的是ResultSet对象(本质是标签列表),不能直接调用find_all方法,必须遍历列表中的每个表格元素单独处理。
修改后的代码
import pandas as pd import requests from bs4 import BeautifulSoup def updatebr(): url='https://wiki.warthunder.com/List_of_vehicle_battle_ratings' headers =[] r = requests.get(url) soup = BeautifulSoup(r.text, 'html.parser') # 获取所有表格,纠正变量名拼写(原vehical改为vehicles) vehicles = soup.find_all('table') # 从第一个表格提取表头 if vehicles: first_table = vehicles[0] for i in first_table.find_all('th'): title = i.text.strip() # 去除多余空格换行 headers.append(title) df = pd.DataFrame(columns=headers) # 遍历所有表格,提取数据行 for table in vehicles: # 跳过表头行,从第2行开始读取数据 for row in table.find_all('tr')[1:]: data = row.find_all('td') row_data = [td.text.strip() for td in data] # 确保数据列数和表头一致再添加 if len(row_data) == len(headers): df.loc[len(df)] = row_data df.to_excel('brlist.xlsx', index=False) # 去掉Excel中的索引列 updatebr()
关键改动说明
- 将
soup.find('table')改为soup.find_all('table')获取全部表格,同时纠正变量名拼写错误(原vehical改为vehicles) - 仅从第一个表格提取一次表头,避免重复
- 遍历每个表格,对每个表格单独调用
find_all('tr')处理数据行 - 添加
strip()去除文本中的多余空格和换行符,优化数据整洁度 - 导出Excel时添加
index=False,避免生成不必要的索引列
内容的提问来源于stack exchange,提问作者Kalween
相关产品推荐
相关产品推荐

