You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取HHS OIG报告标题的Python爬虫生成空CSV问题求助

问题排查:爬取HHS OIG报告标题生成空CSV

我尝试用Python+BeautifulSoup爬取https://oig.hhs.gov/reports-and-publications/all-reports-and-publications/的政府报告标题(例如"Washington Medicaid Fraud Control Unit: 2023 Inspection")并保存到CSV文件,但代码始终生成空文件。我的代码如下:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# URL of the website to scrape
url = 'https://oig.hhs.gov/reports-and-publications/all-reports-and-publications/'

# Send a GET request to the website
response = requests.get(url)

# Check if the request was successful
if response.status_code == 200:
    # Parse the HTML content of the page with BeautifulSoup
    soup = BeautifulSoup(response.content, 'html.parser')

    # Print the structure for debugging purposes
    print(soup.prettify())

    # Find all the report titles
    titles = []
    for title in soup.find_all('a', class_='item-title'):
        titles.append(title.get_text(strip=True))

    # If no titles are found, print a message
    if not titles:
        print("No titles found. The structure of the page might have changed.")

    # Create a DataFrame from the list of titles
    df = pd.DataFrame(titles, columns=['Report Title'])

    # Save the DataFrame to a CSV file
    df.to_csv('report_titles.csv', index=False)
    print("Report titles have been saved to report_titles.csv")
else:
    print(f"Failed to retrieve the webpage. Status code: {response.status_code}")

问题原因分析

空CSV的核心原因是代码中的soup.find_all('a', class_='item-title')没有匹配到任何元素,导致titles列表为空。主要有两种可能:

  1. 页面结构变更:目标网站的报告标题元素标签或class已不再是a.item-title
  2. 动态内容加载:报告列表是通过JavaScript异步渲染的,requests.get()只能获取初始静态HTML,无法拿到动态加载的内容

解决方案

方案1:修正元素选择器(针对结构变更)

先手动检查页面实际结构:打开目标页面,右键「检查」查看报告标题的标签和class。当前页面的报告标题实际是嵌套在h3标签下的a标签,class为report-title。修改代码中的查找逻辑:

# 替换原有的标题查找代码
titles = []
for title in soup.find_all('a', class_='report-title'):
    titles.append(title.get_text(strip=True))

方案2:处理动态加载(针对JS渲染内容)

如果静态HTML中没有报告列表,说明内容是动态加载的,需要用Selenium获取渲染后的页面:

import pandas as pd
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time

url = 'https://oig.hhs.gov/reports-and-publications/all-reports-and-publications/'

# 配置无头Chrome浏览器
chrome_options = Options()
chrome_options.add_argument("--headless=new")
driver = webdriver.Chrome(options=chrome_options)

driver.get(url)
time.sleep(3)  # 等待页面加载完成

# 获取渲染后的HTML
soup = BeautifulSoup(driver.page_source, 'html.parser')
driver.quit()

# 查找报告标题(根据实际结构调整class)
titles = []
for title in soup.find_all('a', class_='report-title'):
    titles.append(title.get_text(strip=True))

if not titles:
    print("未找到标题,页面结构可能已变更,请重新检查元素选择器。")
else:
    df = pd.DataFrame(titles, columns=['Report Title'])
    df.to_csv('report_titles.csv', index=False)
    print("报告标题已保存至report_titles.csv")

调试技巧

  • 打印soup.find_all('a')查看所有<a>标签,确认目标标题的标签属性
  • 打印response.text查看requests获取的原始HTML,判断是否包含报告内容(如果没有则是动态加载问题)

内容的提问来源于stack exchange,提问作者user23471091

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 12:37:28