You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup仅抓取月度板块的NPPES Data Dissemination链接?

解决方法:精准抓取指定板块的目标链接

问题分析

你需要从目标页面的「Full Replacement Monthly NPI File」板块中提取含「NPPES Data Dissemination」文本的<a>链接,但现有代码会同时抓取「Weekly Incremental NPI Files」板块的同类链接,之前的两种尝试均无效:

  • 设置limit=1仅限制单个<a>标签内的匹配次数,无法过滤不同板块的链接;
  • 正则中使用(未转义,导致正则语法错误((是正则的特殊元字符,需转义为\(才能匹配字面量)。

正确实现思路

先精准定位到「Full Replacement Monthly NPI File」板块的容器,再在该容器内部查找目标链接,从根源上排除其他板块的内容。

修改后的完整代码

import re
from bs4 import BeautifulSoup
import requests
import wget

def get_urls(soup):
    urls = []
    # 定位月度板块的标题元素(页面中该标题为<h3>标签)
    monthly_section_title = soup.find('h3', string=re.compile('Full Replacement Monthly NPI File'))
    if monthly_section_title:
        # 获取标题的下一个兄弟列表容器,该容器包含月度板块的所有链接
        monthly_section = monthly_section_title.find_next_sibling('ul')
        # 在板块容器内筛选符合条件的<a>标签
        for a in monthly_section.find_all('a', href=True):
            if re.search('NPPES Data Dissemination', a.get_text()):
                urls.append(a)
    print('done scraping the url...')
    return urls

def download_and_extract(urls):
    for texts in urls:
        text = str(texts)
        file = text[55:99]
        print('zip file :', file)
        zip_link = texts['href']
        print('Downloading %s :' % zip_link)
        slashurl = zip_link.split('/')
        print(slashurl)
        wget.download("https://download.cms.gov/nppes/" + slashurl[1])

r = requests.get('https://download.cms.gov/nppes/NPI_Files.html')
soup = BeautifulSoup(r.content, 'html.parser')
urls = get_urls(soup)
download_and_extract(urls)

代码说明

  1. 精准定位板块:通过标题文本匹配找到月度板块的入口,确保只处理目标板块内的内容;
  2. 锁定板块内容:利用find_next_sibling获取标题对应的链接列表容器,避免跨板块抓取;
  3. 筛选目标链接:在目标容器内直接匹配含指定文本的<a>标签,逻辑更简洁且无正则转义问题;
  4. 兼容性保障:基于页面现有结构定位元素,稳定性优于全局查找。

内容的提问来源于stack exchange,提问作者sherri pytorch

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 07:40:34