使用Python抓取网页表格数据遇阻,求技术解决方案
网页表格抓取问题解决方法
问题分析
你遇到的pd.read_html()找不到表格的问题,大概率是因为:
- 直接请求缺少浏览器标识,服务器返回的内容不包含完整表格
- Windows路径的转义字符导致保存失败
- CSRF令牌可能过期,请求参数无效
修正后的代码
import requests import pandas as pd from bs4 import BeautifulSoup # 目标请求链接(注意CSRF令牌可能过期,若失效需重新从页面获取) url = 'https://www.swf-wc.usace.army.mil/cgi-bin/rcshtml.pl?day1=05&month1=08&year1=22&Submit1=Submit&page=Reports&report=morn&csrf=cbptbpFggR' # 模拟浏览器请求头,避免被识别为爬虫 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 发送请求并设置正确编码 response = requests.get(url, headers=headers) response.encoding = 'utf-8' # 先检查页面是否存在表格元素 soup = BeautifulSoup(response.text, 'html.parser') tables = soup.find_all('table') if tables: # 用pandas解析表格 df_list = pd.read_html(response.text) if df_list: df = df_list[0] print(df) # 保存到Excel,使用原始字符串避免路径转义问题 df.to_excel(r'C:\Users\jeffr\Desktop\code\swf1.xlsx', index=False) else: print("无法识别页面中的表格结构") else: print("页面中未找到任何表格元素")
关键注意事项
- 请求头设置:必须添加
User-Agent,否则服务器可能返回非预期内容 - 路径处理:Windows文件路径要使用原始字符串
r''或双反斜杠\\,防止转义字符报错 - CSRF令牌有效性:如果当前令牌过期,需要先访问网站首页,从页面源码中提取新的
csrf参数,替换到请求URL中 - 编码处理:设置
response.encoding = 'utf-8'确保页面内容编码正确,避免乱码
内容的提问来源于stack exchange,提问作者Jeff Mlak
相关产品推荐
相关产品推荐

