You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+BS4爬虫CSV输出异常求助:错误条目与目标条目缺失

解决方案

问题原因分析

  1. 错误首个条目:代码抓取了页面中指向自身的导航链接(文本为"Resolution Agreements",URL为当前页面路径),该链接符合你设置的路径包含条件,因此被写入CSV。
  2. 遗漏目标条目:目标条目链接的路径为/hipaa/news/...,并不包含你设置的/agreements/路径,被过滤条件排除;同时代码抓取了页面所有<a>标签,范围过广,没有聚焦到目标条目所在的列表容器。

修改后的代码

import requests
from bs4 import BeautifulSoup
import csv

# URL of the page to scrape
url = "https://www.hhs.gov/hipaa/for-professionals/compliance-enforcement/agreements/index.html"

# Send a GET request to the URL
response = requests.get(url)

# Parse the HTML content
soup = BeautifulSoup(response.text, 'html.parser')

# 定位到包含所有决议协议条目的容器,仅抓取该容器内的链接
agreements_container = soup.find('div', class_='block--list')
links = agreements_container.find_all('a') if agreements_container else []

# Create and open a CSV file in write mode
with open('hipaa_links.csv', mode='w', newline='', encoding='utf-8') as file:
    # Create a CSV writer object
    writer = csv.writer(file)

    # Write the header row
    writer.writerow(['Title', 'URL'])

    # Iterate over each link
    for link in links:
        # Extract link URL
        link_url = link.get('href')

        # Extract link title
        link_title = link.text.strip()

        # 仅保留非空的标题和URL,容器内的链接均为目标条目
        if link_url and link_title:
            # 可选:拼接为绝对URL,方便直接访问
            full_url = f"https://www.hhs.gov{link_url}" if link_url.startswith('/') else link_url
            # Write the link title and URL to the CSV file
            writer.writerow([link_title, full_url])

print("Data has been written to hipaa_links.csv")

关键修改点

  • 缩小抓取范围:通过div.block--list定位到目标条目所在的容器,避免抓取导航、页脚等无关链接。
  • 移除错误路径过滤:原条件过滤掉了/news/路径的目标条目,现在基于容器抓取,无需路径检查。
  • 优化URL处理:可选拼接绝对URL,直接生成可访问的完整链接(保留相对路径也可正常使用)。

内容的提问来源于stack exchange,提问作者user23471091

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 22:07:44