如何用BeautifulSoup通过部分文本匹配提取指定下载URL?
解决方法
问题根源
你用string参数匹配失败,是因为string只会匹配当前标签直接包含的文本,而目标文件名(比如sentinel-3.2022335...CIcyano...LakeOkee.tif)大概率在section的子标签(如<a>、<p>)里,不是section标签的直接文本内容,所以匹配不到。
可行代码方案
方案1:通过Lambda检查子元素文本
直接在查找section时,检查其内部是否存在包含目标字符串的文本节点:
from bs4 import BeautifulSoup import re # 假设soup是已解析的页面对象 target_section = soup.find( 'section', class_='onecol habonecol', lambda tag: tag.find(text=re.compile(r'CIcyano')) is not None ) if target_section: # 提取下载链接 download_url = target_section.find('a', href=True)['href'] print(download_url)
方案2:遍历所有目标section,检查完整文本内容
先获取所有符合class的section,再逐一检查其下所有文本是否包含目标字符串:
from bs4 import BeautifulSoup sections = soup.find_all('section', class_='onecol habonecol') for section in sections: # 获取section下所有子元素的文本并拼接 full_content = section.get_text(strip=True) if 'CIcyano' in full_content: download_url = section.find('a', href=True)['href'] print(download_url) break # 找到第一个匹配项后停止遍历
方案3:直接定位包含目标文件名的链接,再找父section
如果确定文件名在<a>标签文本里,可直接定位链接再回溯到section:
from bs4 import BeautifulSoup import re target_link = soup.find('a', href=True, text=re.compile(r'CIcyano')) if target_link: target_section = target_link.find_parent('section', class_='onecol habonecol') download_url = target_link['href'] print(download_url)
内容的提问来源于stack exchange,提问作者Koelker12
相关产品推荐
相关产品推荐

