在Spyder IDE中用Python BeautifulSoup爬取表格遇数据拼接问题求助
爬取NOAA国际飓风表格的问题解决指南
问题说明
在Spyder IDE中爬取https://www.aoml.noaa.gov/hrd/hurdat/International_Hurricanes.html的表格时,现有代码出现以下问题:
- 表头提取未收集到有效数据
- 行数据拼接错误,出现“第1行+所有行”“第2行+所有行”的异常情况
- 无法将爬取结果保存为CSV文件
原代码如下:
import os print(os.getcwd()) import pandas as pd import requests from bs4 import BeautifulSoup import csv url = 'https://www.aoml.noaa.gov/hrd/hurdat/International_Hurricanes.html' response = requests.get(url) html_content = response.text soup = BeautifulSoup(html_content, 'html.parser') table = soup.find('table', {'class': 'content'}) data = table.find_all('tr') headers = [] for header in data[2].find_all('td'): header_lines = header.text.strip().split('\r\n') print(headers) data = [] for row in rows.find_all('tr')[3:]:`find data after 3rd row` cols = row.find_all('td') cols = [col.get_text(strip=True) for col in cols] if cols: data.append(cols)
问题分析与修正方案
原代码核心问题
- 表头提取时,仅拆分文本但未将结果添加到
headers列表 - 使用了未定义的
rows变量,应该调用已定位的table对象 - 缩进错误:
if cols:语句在循环外,导致仅最后一行数据被添加 - 行索引定位错误,未准确找到表头和数据行的起始位置
修正后的完整代码
import requests from bs4 import BeautifulSoup import csv # 目标URL url = 'https://www.aoml.noaa.gov/hrd/hurdat/International_Hurricanes.html' # 获取网页内容并设置编码 response = requests.get(url) response.encoding = 'utf-8' soup = BeautifulSoup(response.text, 'html.parser') # 定位目标表格 table = soup.find('table', {'class': 'content'}) if not table: print("未找到目标表格,请检查网页结构是否变更") exit() # 提取表头:第2个<tr>是表头行(索引从0开始) headers = [] header_row = table.find_all('tr')[1] for cell in header_row.find_all('td'): # 处理表头内的换行,取第一行作为表头名称 header_text = cell.text.strip().split('\n')[0].strip() headers.append(header_text) # 提取数据行:从第3个<tr>开始(索引2) data_rows = [] for row in table.find_all('tr')[2:]: cells = row.find_all('td') # 校验单元格数量与表头一致,过滤无效行 if len(cells) == len(headers): row_data = [cell.text.strip() for cell in cells] data_rows.append(row_data) # 保存为CSV文件 output_file = 'international_hurricanes.csv' with open(output_file, 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(headers) # 写入表头 writer.writerows(data_rows) # 写入所有数据行 print(f"数据已成功保存到 {output_file}")
关键修正说明
- 编码设置:添加
response.encoding = 'utf-8'避免特殊字符乱码 - 表头处理:准确找到表头所在行,拆分并提取有效表头文本
- 变量与缩进:替换未定义变量,将数据判断逻辑放入循环内,确保每行数据都被正确收集
- 数据校验:通过单元格数量匹配表头,过滤空行或格式异常的行
- CSV保存:使用
csv模块规范写入,指定newline=''避免生成多余空行
内容的提问来源于stack exchange,提问作者Rimi
相关产品推荐
相关产品推荐

