Python+BS4爬虫CSV输出异常求助:错误条目与目标条目缺失
解决方案
问题原因分析
- 错误首个条目:代码抓取了页面中指向自身的导航链接(文本为"Resolution Agreements",URL为当前页面路径),该链接符合你设置的路径包含条件,因此被写入CSV。
- 遗漏目标条目:目标条目链接的路径为
/hipaa/news/...,并不包含你设置的/agreements/路径,被过滤条件排除;同时代码抓取了页面所有<a>标签,范围过广,没有聚焦到目标条目所在的列表容器。
修改后的代码
import requests from bs4 import BeautifulSoup import csv # URL of the page to scrape url = "https://www.hhs.gov/hipaa/for-professionals/compliance-enforcement/agreements/index.html" # Send a GET request to the URL response = requests.get(url) # Parse the HTML content soup = BeautifulSoup(response.text, 'html.parser') # 定位到包含所有决议协议条目的容器,仅抓取该容器内的链接 agreements_container = soup.find('div', class_='block--list') links = agreements_container.find_all('a') if agreements_container else [] # Create and open a CSV file in write mode with open('hipaa_links.csv', mode='w', newline='', encoding='utf-8') as file: # Create a CSV writer object writer = csv.writer(file) # Write the header row writer.writerow(['Title', 'URL']) # Iterate over each link for link in links: # Extract link URL link_url = link.get('href') # Extract link title link_title = link.text.strip() # 仅保留非空的标题和URL,容器内的链接均为目标条目 if link_url and link_title: # 可选:拼接为绝对URL,方便直接访问 full_url = f"https://www.hhs.gov{link_url}" if link_url.startswith('/') else link_url # Write the link title and URL to the CSV file writer.writerow([link_title, full_url]) print("Data has been written to hipaa_links.csv")
关键修改点
- 缩小抓取范围:通过
div.block--list定位到目标条目所在的容器,避免抓取导航、页脚等无关链接。 - 移除错误路径过滤:原条件过滤掉了
/news/路径的目标条目,现在基于容器抓取,无需路径检查。 - 优化URL处理:可选拼接绝对URL,直接生成可访问的完整链接(保留相对路径也可正常使用)。
内容的提问来源于stack exchange,提问作者user23471091
相关产品推荐
相关产品推荐

