爬取HHS OIG报告标题的Python爬虫生成空CSV问题求助
问题排查:爬取HHS OIG报告标题生成空CSV
我尝试用Python+BeautifulSoup爬取https://oig.hhs.gov/reports-and-publications/all-reports-and-publications/的政府报告标题(例如"Washington Medicaid Fraud Control Unit: 2023 Inspection")并保存到CSV文件,但代码始终生成空文件。我的代码如下:
import requests from bs4 import BeautifulSoup import pandas as pd # URL of the website to scrape url = 'https://oig.hhs.gov/reports-and-publications/all-reports-and-publications/' # Send a GET request to the website response = requests.get(url) # Check if the request was successful if response.status_code == 200: # Parse the HTML content of the page with BeautifulSoup soup = BeautifulSoup(response.content, 'html.parser') # Print the structure for debugging purposes print(soup.prettify()) # Find all the report titles titles = [] for title in soup.find_all('a', class_='item-title'): titles.append(title.get_text(strip=True)) # If no titles are found, print a message if not titles: print("No titles found. The structure of the page might have changed.") # Create a DataFrame from the list of titles df = pd.DataFrame(titles, columns=['Report Title']) # Save the DataFrame to a CSV file df.to_csv('report_titles.csv', index=False) print("Report titles have been saved to report_titles.csv") else: print(f"Failed to retrieve the webpage. Status code: {response.status_code}")
问题原因分析
空CSV的核心原因是代码中的soup.find_all('a', class_='item-title')没有匹配到任何元素,导致titles列表为空。主要有两种可能:
- 页面结构变更:目标网站的报告标题元素标签或class已不再是
a.item-title - 动态内容加载:报告列表是通过JavaScript异步渲染的,
requests.get()只能获取初始静态HTML,无法拿到动态加载的内容
解决方案
方案1:修正元素选择器(针对结构变更)
先手动检查页面实际结构:打开目标页面,右键「检查」查看报告标题的标签和class。当前页面的报告标题实际是嵌套在h3标签下的a标签,class为report-title。修改代码中的查找逻辑:
# 替换原有的标题查找代码 titles = [] for title in soup.find_all('a', class_='report-title'): titles.append(title.get_text(strip=True))
方案2:处理动态加载(针对JS渲染内容)
如果静态HTML中没有报告列表,说明内容是动态加载的,需要用Selenium获取渲染后的页面:
import pandas as pd from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options import time url = 'https://oig.hhs.gov/reports-and-publications/all-reports-and-publications/' # 配置无头Chrome浏览器 chrome_options = Options() chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) driver.get(url) time.sleep(3) # 等待页面加载完成 # 获取渲染后的HTML soup = BeautifulSoup(driver.page_source, 'html.parser') driver.quit() # 查找报告标题(根据实际结构调整class) titles = [] for title in soup.find_all('a', class_='report-title'): titles.append(title.get_text(strip=True)) if not titles: print("未找到标题,页面结构可能已变更,请重新检查元素选择器。") else: df = pd.DataFrame(titles, columns=['Report Title']) df.to_csv('report_titles.csv', index=False) print("报告标题已保存至report_titles.csv")
调试技巧
- 打印
soup.find_all('a')查看所有<a>标签,确认目标标题的标签属性 - 打印
response.text查看requests获取的原始HTML,判断是否包含报告内容(如果没有则是动态加载问题)
内容的提问来源于stack exchange,提问作者user23471091
相关产品推荐
相关产品推荐

