如何去除网页抓取数据中的政府官网头部提示内容?
解决方案:移除HHS.gov子页面开头的固定政府提示内容
核心思路
HHS.gov子页面开头的固定提示属于标准化免责声明,可通过精准匹配文本内容或定位HTML结构排除两种方式移除,后者更稳定(避免文本细微变动导致失效)。
方法1:基于固定文本精准移除
先抓取任意一个子页面,复制开头的完整提示文本(比如类似"U.S. Department of Health and Human Services (HHS) provides this information for educational purposes only..."),然后在代码中对抓取到的描述文本做针对性截取:
修改代码中写入CSV前的逻辑:
# 抓取子页面文本 description = scrape_page_text(link_url) # 替换为实际抓取到的完整固定提示文本 fixed_disclaimer = "U.S. Department of Health and Human Services (HHS) provides this information for educational purposes only. It is not legal advice or a legal determination of any kind." # 检查描述是否以提示开头,若是则移除 if description.startswith(fixed_disclaimer): description = description[len(fixed_disclaimer):].strip()
方法2:基于HTML结构排除提示(推荐)
政府网站的免责声明通常放在特定HTML容器中,可直接定位并移除该容器后再抓取正文:
修改scrape_page_text函数:
def scrape_page_text(url): response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # 定位并移除免责声明容器(根据实际页面结构调整选择器) # 示例:若提示在class为"gov-disclaimer"的p标签中 disclaimer = soup.find('p', class_='gov-disclaimer') # 或者若提示是页面前2个p标签,直接跳过 # paragraphs = soup.find_all('p')[2:] if disclaimer: disclaimer.decompose() # 从DOM中移除该标签 # 抓取剩余段落文本 paragraphs = soup.find_all('p') text = ' '.join([p.get_text().strip() for p in paragraphs]) return text.strip()
如何确定HTML选择器?
打开任意子页面,右键提示文本→检查元素,查看该文本所在标签的class/id或位置,调整代码中的选择器即可。
完整修改后的代码
import requests from bs4 import BeautifulSoup import csv from urllib.parse import urljoin # Function to scrape text from a given URL, excluding disclaimer def scrape_page_text(url): response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # 移除开头的免责声明(示例:假设提示在第一个p标签,或根据实际结构调整) # 若提示是特定class,替换为soup.find('p', class_='your-disclaimer-class') first_p = soup.find('p') # 可添加文本判断确保是免责声明:比如判断是否包含"U.S. Department of Health and Human Services" if first_p and "U.S. Department of Health and Human Services" in first_p.get_text(): first_p.decompose() paragraphs = soup.find_all('p') text = ' '.join([p.get_text().strip() for p in paragraphs]) return text.strip() # Base URL base_url = "https://www.hhs.gov/hipaa/for-professionals/compliance-enforcement/agreements/" # URL of the page to scrape url = urljoin(base_url, "index.html") response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') # Find the innermost l-content div content_divs = soup.find_all('div', class_='l-content') content_div = content_divs[-1] links = content_div.find_all('a') with open('hipaa_links.csv', mode='w', newline='\n', encoding='utf-8') as file: writer = csv.writer(file) writer.writerow(['Title', 'URL', 'Description']) for link in links: link_url = urljoin(base_url, link.get('href')) link_title = link.text.strip() description = scrape_page_text(link_url) if link_url and link_title: writer.writerow([link_title, link_url, description]) print("Data has been written to hipaa_links.csv")
内容的提问来源于stack exchange,提问作者user23471091
相关产品推荐
相关产品推荐

