You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python抓取网页表格数据遇阻,求技术解决方案

网页表格抓取问题解决方法

问题分析

你遇到的pd.read_html()找不到表格的问题,大概率是因为:

  • 直接请求缺少浏览器标识,服务器返回的内容不包含完整表格
  • Windows路径的转义字符导致保存失败
  • CSRF令牌可能过期,请求参数无效

修正后的代码

import requests
import pandas as pd
from bs4 import BeautifulSoup

# 目标请求链接(注意CSRF令牌可能过期,若失效需重新从页面获取)
url = 'https://www.swf-wc.usace.army.mil/cgi-bin/rcshtml.pl?day1=05&month1=08&year1=22&Submit1=Submit&page=Reports&report=morn&csrf=cbptbpFggR'

# 模拟浏览器请求头,避免被识别为爬虫
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

# 发送请求并设置正确编码
response = requests.get(url, headers=headers)
response.encoding = 'utf-8'

# 先检查页面是否存在表格元素
soup = BeautifulSoup(response.text, 'html.parser')
tables = soup.find_all('table')

if tables:
    # 用pandas解析表格
    df_list = pd.read_html(response.text)
    if df_list:
        df = df_list[0]
        print(df)
        # 保存到Excel,使用原始字符串避免路径转义问题
        df.to_excel(r'C:\Users\jeffr\Desktop\code\swf1.xlsx', index=False)
    else:
        print("无法识别页面中的表格结构")
else:
    print("页面中未找到任何表格元素")

关键注意事项

  • 请求头设置:必须添加User-Agent,否则服务器可能返回非预期内容
  • 路径处理:Windows文件路径要使用原始字符串r''或双反斜杠\\,防止转义字符报错
  • CSRF令牌有效性:如果当前令牌过期,需要先访问网站首页,从页面源码中提取新的csrf参数,替换到请求URL中
  • 编码处理:设置response.encoding = 'utf-8'确保页面内容编码正确,避免乱码

内容的提问来源于stack exchange,提问作者Jeff Mlak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 02:06:52