爬取HHS OIG报告网站生成空CSV,求代码修改建议
解决HHS OIG报告爬取后CSV为空的问题
尝试爬取HHS OIG报告页面的报告信息(标题、审计项、机构、日期),但执行代码后生成的CSV文件为空,核心问题在于请求拦截和元素选择器错误,以下是修改方案:
问题根源
- 无浏览器标识被拦截:服务器会拦截未携带
User-Agent的请求,返回的内容不含目标数据 - 元素选择器不匹配:原代码使用
div.media定位报告容器,但当前页面的报告项实际使用div.views-row作为容器,导致无法抓取到任何条目
修改后的完整代码
import requests from bs4 import BeautifulSoup import csv url = "https://oig.hhs.gov/reports-and-publications/all-reports-and-publications/" # 模拟浏览器请求头,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) # 检查请求是否成功,失败则抛出异常 response.raise_for_status() soup = BeautifulSoup(response.content, "html.parser") # 匹配页面实际的报告条目容器 reports = soup.find_all("div", class_="views-row") report_data = [] for report in reports: # 提取标题,增加空值判断 title_elem = report.find("h3", class_="report-title") title = title_elem.get_text(strip=True) if title_elem else "N/A" # 提取元数据容器,统一处理审计项、机构、日期 meta_container = report.find("div", class_="report-meta") audit = meta_container.find("span", class_="audit").get_text(strip=True) if meta_container and meta_container.find("span", class_="audit") else "N/A" agency = meta_container.find("span", class_="agency").get_text(strip=True) if meta_container and meta_container.find("span", class_="agency") else "N/A" date = meta_container.find("span", class_="date").get_text(strip=True) if meta_container and meta_container.find("span", class_="date") else "N/A" report_data.append({ "Title": title, "Audit": audit, "Agency": agency, "Date": date }) # 导出CSV csv_file = "reports_data.csv" with open(csv_file, mode='w', newline='', encoding='utf-8') as file: writer = csv.DictWriter(file, fieldnames=["Title", "Audit", "Agency", "Date"]) writer.writeheader() for data in report_data: writer.writerow(data) print(f"Data exported to {csv_file}")
关键修改点说明
- 添加请求头:通过
User-Agent模拟浏览器访问,确保服务器返回完整页面内容 - 修正容器选择器:将
div.media替换为div.views-row,匹配页面实际的报告条目结构 - 增强容错性:对所有元素提取增加空值判断,避免因个别条目结构异常导致代码中断
- 增加请求校验:
response.raise_for_status()会在请求失败(如403、500)时抛出异常,便于快速排查网络问题
内容的提问来源于stack exchange,提问作者user23471091
相关产品推荐
相关产品推荐

