You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何去除网页抓取数据中的政府官网头部提示内容?

解决方案:移除HHS.gov子页面开头的固定政府提示内容

核心思路

HHS.gov子页面开头的固定提示属于标准化免责声明,可通过精准匹配文本内容或定位HTML结构排除两种方式移除,后者更稳定(避免文本细微变动导致失效)。


方法1:基于固定文本精准移除

先抓取任意一个子页面,复制开头的完整提示文本(比如类似"U.S. Department of Health and Human Services (HHS) provides this information for educational purposes only..."),然后在代码中对抓取到的描述文本做针对性截取:

修改代码中写入CSV前的逻辑:

# 抓取子页面文本
description = scrape_page_text(link_url)

# 替换为实际抓取到的完整固定提示文本
fixed_disclaimer = "U.S. Department of Health and Human Services (HHS) provides this information for educational purposes only. It is not legal advice or a legal determination of any kind."
# 检查描述是否以提示开头,若是则移除
if description.startswith(fixed_disclaimer):
    description = description[len(fixed_disclaimer):].strip()

方法2:基于HTML结构排除提示(推荐)

政府网站的免责声明通常放在特定HTML容器中,可直接定位并移除该容器后再抓取正文:

修改scrape_page_text函数:

def scrape_page_text(url):
    response = requests.get(url)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # 定位并移除免责声明容器(根据实际页面结构调整选择器)
    # 示例:若提示在class为"gov-disclaimer"的p标签中
    disclaimer = soup.find('p', class_='gov-disclaimer')
    # 或者若提示是页面前2个p标签,直接跳过
    # paragraphs = soup.find_all('p')[2:]
    
    if disclaimer:
        disclaimer.decompose()  # 从DOM中移除该标签
    
    # 抓取剩余段落文本
    paragraphs = soup.find_all('p')
    text = ' '.join([p.get_text().strip() for p in paragraphs])
    return text.strip()

如何确定HTML选择器?

打开任意子页面,右键提示文本→检查元素,查看该文本所在标签的class/id或位置,调整代码中的选择器即可。


完整修改后的代码

import requests
from bs4 import BeautifulSoup
import csv
from urllib.parse import urljoin

# Function to scrape text from a given URL, excluding disclaimer
def scrape_page_text(url):
    response = requests.get(url)
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # 移除开头的免责声明(示例:假设提示在第一个p标签,或根据实际结构调整)
    # 若提示是特定class,替换为soup.find('p', class_='your-disclaimer-class')
    first_p = soup.find('p')
    # 可添加文本判断确保是免责声明:比如判断是否包含"U.S. Department of Health and Human Services"
    if first_p and "U.S. Department of Health and Human Services" in first_p.get_text():
        first_p.decompose()
    
    paragraphs = soup.find_all('p')
    text = ' '.join([p.get_text().strip() for p in paragraphs])
    return text.strip()

# Base URL
base_url = "https://www.hhs.gov/hipaa/for-professionals/compliance-enforcement/agreements/"

# URL of the page to scrape
url = urljoin(base_url, "index.html")

response = requests.get(url)
soup = BeautifulSoup(response.text, 'html.parser')

# Find the innermost l-content div
content_divs = soup.find_all('div', class_='l-content')
content_div = content_divs[-1]

links = content_div.find_all('a')

with open('hipaa_links.csv', mode='w', newline='\n', encoding='utf-8') as file:
    writer = csv.writer(file)
    writer.writerow(['Title', 'URL', 'Description'])

    for link in links:
        link_url = urljoin(base_url, link.get('href'))
        link_title = link.text.strip()
        description = scrape_page_text(link_url)

        if link_url and link_title:
            writer.writerow([link_title, link_url, description])

print("Data has been written to hipaa_links.csv")

内容的提问来源于stack exchange,提问作者user23471091

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 22:34:51