使用Python与BeautifulSoup爬取网站表格返回None或空的解决方法
解决爬取网页表格返回空的问题
问题根源
直接使用requests.get()请求目标网站时,服务器会识别出非浏览器的爬虫请求,返回的页面不包含目标表格内容,导致BeautifulSoup无法定位到表格元素。
解决方案
1. 添加请求头伪装浏览器访问
给请求添加User-Agent头,模拟正常浏览器的访问行为,获取完整的页面内容:
import requests import pandas as pd from bs4 import BeautifulSoup url = "https://aviation-safety.net/database/year/2024/1" # 模拟Chrome浏览器的请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers) # 确认请求成功(状态码为200) print(response.status_code) soup = BeautifulSoup(response.text, "lxml") table = soup.find("table", class_ = "hp") print(table)
2. 用pandas直接读取表格(更高效)
既然最终目标是生成CSV文件,可直接使用pandas.read_html()自动解析页面中的表格,无需手动用BeautifulSoup定位:
import pandas as pd url = "https://aviation-safety.net/database/year/2024/1" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } # 读取页面中所有表格,返回DataFrame列表 dfs = pd.read_html(url, headers=headers) # 目标表格为页面中的第一个表格,直接提取 target_df = dfs[0] # 导出为CSV文件 target_df.to_csv("aviation_safety_2024.csv", index=False, encoding="utf-8")
验证说明
添加请求头后,服务器会返回包含完整表格的页面,此时无论是用BeautifulSoup查找还是pandas.read_html都能正常获取数据。pandas.read_html会自动处理表格结构,直接生成DataFrame,后续导出CSV的步骤更简洁。
内容的提问来源于stack exchange,提问作者Douglas OB
相关产品推荐
相关产品推荐

