You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Spyder IDE中用Python BeautifulSoup爬取表格遇数据拼接问题求助

爬取NOAA国际飓风表格的问题解决指南

问题说明

在Spyder IDE中爬取https://www.aoml.noaa.gov/hrd/hurdat/International_Hurricanes.html的表格时,现有代码出现以下问题:

  • 表头提取未收集到有效数据
  • 行数据拼接错误,出现“第1行+所有行”“第2行+所有行”的异常情况
  • 无法将爬取结果保存为CSV文件

原代码如下:

import os
print(os.getcwd())

import pandas as pd
import requests  
from bs4 import BeautifulSoup  
import csv  

url = 'https://www.aoml.noaa.gov/hrd/hurdat/International_Hurricanes.html'
response = requests.get(url)  
html_content = response.text  
soup = BeautifulSoup(html_content, 'html.parser')  

table = soup.find('table', {'class': 'content'})
data = table.find_all('tr')

headers = []
for header in data[2].find_all('td'): 
    header_lines = header.text.strip().split('\r\n')

print(headers)

data = []
for row in rows.find_all('tr')[3:]:`find data after 3rd row`
    cols = row.find_all('td')
    cols = [col.get_text(strip=True) for col in cols]
if cols:
    data.append(cols)

问题分析与修正方案

原代码核心问题

  1. 表头提取时,仅拆分文本但未将结果添加到headers列表
  2. 使用了未定义的rows变量,应该调用已定位的table对象
  3. 缩进错误:if cols:语句在循环外,导致仅最后一行数据被添加
  4. 行索引定位错误,未准确找到表头和数据行的起始位置

修正后的完整代码

import requests
from bs4 import BeautifulSoup
import csv

# 目标URL
url = 'https://www.aoml.noaa.gov/hrd/hurdat/International_Hurricanes.html'

# 获取网页内容并设置编码
response = requests.get(url)
response.encoding = 'utf-8'
soup = BeautifulSoup(response.text, 'html.parser')

# 定位目标表格
table = soup.find('table', {'class': 'content'})
if not table:
    print("未找到目标表格,请检查网页结构是否变更")
    exit()

# 提取表头:第2个<tr>是表头行(索引从0开始)
headers = []
header_row = table.find_all('tr')[1]
for cell in header_row.find_all('td'):
    # 处理表头内的换行,取第一行作为表头名称
    header_text = cell.text.strip().split('\n')[0].strip()
    headers.append(header_text)

# 提取数据行:从第3个<tr>开始(索引2)
data_rows = []
for row in table.find_all('tr')[2:]:
    cells = row.find_all('td')
    # 校验单元格数量与表头一致,过滤无效行
    if len(cells) == len(headers):
        row_data = [cell.text.strip() for cell in cells]
        data_rows.append(row_data)

# 保存为CSV文件
output_file = 'international_hurricanes.csv'
with open(output_file, 'w', newline='', encoding='utf-8') as f:
    writer = csv.writer(f)
    writer.writerow(headers)  # 写入表头
    writer.writerows(data_rows)  # 写入所有数据行

print(f"数据已成功保存到 {output_file}")

关键修正说明

  • 编码设置:添加response.encoding = 'utf-8'避免特殊字符乱码
  • 表头处理:准确找到表头所在行,拆分并提取有效表头文本
  • 变量与缩进:替换未定义变量,将数据判断逻辑放入循环内,确保每行数据都被正确收集
  • 数据校验:通过单元格数量匹配表头,过滤空行或格式异常的行
  • CSV保存:使用csv模块规范写入,指定newline=''避免生成多余空行

内容的提问来源于stack exchange,提问作者Rimi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 23:30:56