如何用BeautifulSoup仅抓取月度板块的NPPES Data Dissemination链接?
解决方法:精准抓取指定板块的目标链接
问题分析
你需要从目标页面的「Full Replacement Monthly NPI File」板块中提取含「NPPES Data Dissemination」文本的<a>链接,但现有代码会同时抓取「Weekly Incremental NPI Files」板块的同类链接,之前的两种尝试均无效:
- 设置
limit=1仅限制单个<a>标签内的匹配次数,无法过滤不同板块的链接; - 正则中使用
(未转义,导致正则语法错误((是正则的特殊元字符,需转义为\(才能匹配字面量)。
正确实现思路
先精准定位到「Full Replacement Monthly NPI File」板块的容器,再在该容器内部查找目标链接,从根源上排除其他板块的内容。
修改后的完整代码
import re from bs4 import BeautifulSoup import requests import wget def get_urls(soup): urls = [] # 定位月度板块的标题元素(页面中该标题为<h3>标签) monthly_section_title = soup.find('h3', string=re.compile('Full Replacement Monthly NPI File')) if monthly_section_title: # 获取标题的下一个兄弟列表容器,该容器包含月度板块的所有链接 monthly_section = monthly_section_title.find_next_sibling('ul') # 在板块容器内筛选符合条件的<a>标签 for a in monthly_section.find_all('a', href=True): if re.search('NPPES Data Dissemination', a.get_text()): urls.append(a) print('done scraping the url...') return urls def download_and_extract(urls): for texts in urls: text = str(texts) file = text[55:99] print('zip file :', file) zip_link = texts['href'] print('Downloading %s :' % zip_link) slashurl = zip_link.split('/') print(slashurl) wget.download("https://download.cms.gov/nppes/" + slashurl[1]) r = requests.get('https://download.cms.gov/nppes/NPI_Files.html') soup = BeautifulSoup(r.content, 'html.parser') urls = get_urls(soup) download_and_extract(urls)
代码说明
- 精准定位板块:通过标题文本匹配找到月度板块的入口,确保只处理目标板块内的内容;
- 锁定板块内容:利用
find_next_sibling获取标题对应的链接列表容器,避免跨板块抓取; - 筛选目标链接:在目标容器内直接匹配含指定文本的
<a>标签,逻辑更简洁且无正则转义问题; - 兼容性保障:基于页面现有结构定位元素,稳定性优于全局查找。
内容的提问来源于stack exchange,提问作者sherri pytorch
相关产品推荐
相关产品推荐

