You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取HHS OIG报告网站生成空CSV,求代码修改建议

解决HHS OIG报告爬取后CSV为空的问题

尝试爬取HHS OIG报告页面的报告信息(标题、审计项、机构、日期),但执行代码后生成的CSV文件为空,核心问题在于请求拦截和元素选择器错误,以下是修改方案:

问题根源

  1. 无浏览器标识被拦截:服务器会拦截未携带User-Agent的请求,返回的内容不含目标数据
  2. 元素选择器不匹配:原代码使用div.media定位报告容器,但当前页面的报告项实际使用div.views-row作为容器,导致无法抓取到任何条目

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import csv

url = "https://oig.hhs.gov/reports-and-publications/all-reports-and-publications/"
# 模拟浏览器请求头,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}
response = requests.get(url, headers=headers)
# 检查请求是否成功,失败则抛出异常
response.raise_for_status()
soup = BeautifulSoup(response.content, "html.parser")

# 匹配页面实际的报告条目容器
reports = soup.find_all("div", class_="views-row")

report_data = []

for report in reports:
    # 提取标题,增加空值判断
    title_elem = report.find("h3", class_="report-title")
    title = title_elem.get_text(strip=True) if title_elem else "N/A"
    
    # 提取元数据容器,统一处理审计项、机构、日期
    meta_container = report.find("div", class_="report-meta")
    audit = meta_container.find("span", class_="audit").get_text(strip=True) if meta_container and meta_container.find("span", class_="audit") else "N/A"
    agency = meta_container.find("span", class_="agency").get_text(strip=True) if meta_container and meta_container.find("span", class_="agency") else "N/A"
    date = meta_container.find("span", class_="date").get_text(strip=True) if meta_container and meta_container.find("span", class_="date") else "N/A"
    
    report_data.append({
        "Title": title,
        "Audit": audit,
        "Agency": agency,
        "Date": date
    })

# 导出CSV
csv_file = "reports_data.csv"
with open(csv_file, mode='w', newline='', encoding='utf-8') as file:
    writer = csv.DictWriter(file, fieldnames=["Title", "Audit", "Agency", "Date"])
    writer.writeheader()
    for data in report_data:
        writer.writerow(data)

print(f"Data exported to {csv_file}")

关键修改点说明

  • 添加请求头:通过User-Agent模拟浏览器访问,确保服务器返回完整页面内容
  • 修正容器选择器:将div.media替换为div.views-row,匹配页面实际的报告条目结构
  • 增强容错性:对所有元素提取增加空值判断,避免因个别条目结构异常导致代码中断
  • 增加请求校验:response.raise_for_status()会在请求失败(如403、500)时抛出异常,便于快速排查网络问题

内容的提问来源于stack exchange,提问作者user23471091

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 09:35:07