You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python与BeautifulSoup爬取网站表格返回None或空的解决方法

解决爬取网页表格返回空的问题

问题根源

直接使用requests.get()请求目标网站时,服务器会识别出非浏览器的爬虫请求,返回的页面不包含目标表格内容,导致BeautifulSoup无法定位到表格元素。

解决方案

1. 添加请求头伪装浏览器访问

给请求添加User-Agent头,模拟正常浏览器的访问行为,获取完整的页面内容:

import requests
import pandas as pd 
from bs4 import BeautifulSoup

url = "https://aviation-safety.net/database/year/2024/1"
# 模拟Chrome浏览器的请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
}
response = requests.get(url, headers=headers)

# 确认请求成功(状态码为200)
print(response.status_code)

soup = BeautifulSoup(response.text, "lxml")
table = soup.find("table", class_ = "hp")
print(table)

2. 用pandas直接读取表格(更高效)

既然最终目标是生成CSV文件,可直接使用pandas.read_html()自动解析页面中的表格,无需手动用BeautifulSoup定位:

import pandas as pd

url = "https://aviation-safety.net/database/year/2024/1"
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
}
# 读取页面中所有表格,返回DataFrame列表
dfs = pd.read_html(url, headers=headers)
# 目标表格为页面中的第一个表格,直接提取
target_df = dfs[0]
# 导出为CSV文件
target_df.to_csv("aviation_safety_2024.csv", index=False, encoding="utf-8")

验证说明

添加请求头后,服务器会返回包含完整表格的页面,此时无论是用BeautifulSoup查找还是pandas.read_html都能正常获取数据。pandas.read_html会自动处理表格结构,直接生成DataFrame,后续导出CSV的步骤更简洁。

内容的提问来源于stack exchange,提问作者Douglas OB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 22:45:03