You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup通过部分文本匹配提取指定下载URL?

解决方法

问题根源

你用string参数匹配失败,是因为string只会匹配当前标签直接包含的文本,而目标文件名(比如sentinel-3.2022335...CIcyano...LakeOkee.tif)大概率在section的子标签(如<a>、<p>)里,不是section标签的直接文本内容,所以匹配不到。

可行代码方案

方案1:通过Lambda检查子元素文本

直接在查找section时,检查其内部是否存在包含目标字符串的文本节点:

from bs4 import BeautifulSoup
import re

# 假设soup是已解析的页面对象
target_section = soup.find(
    'section',
    class_='onecol habonecol',
    lambda tag: tag.find(text=re.compile(r'CIcyano')) is not None
)

if target_section:
    # 提取下载链接
    download_url = target_section.find('a', href=True)['href']
    print(download_url)

方案2:遍历所有目标section,检查完整文本内容

先获取所有符合class的section,再逐一检查其下所有文本是否包含目标字符串:

from bs4 import BeautifulSoup

sections = soup.find_all('section', class_='onecol habonecol')
for section in sections:
    # 获取section下所有子元素的文本并拼接
    full_content = section.get_text(strip=True)
    if 'CIcyano' in full_content:
        download_url = section.find('a', href=True)['href']
        print(download_url)
        break  # 找到第一个匹配项后停止遍历

方案3:直接定位包含目标文件名的链接,再找父section

如果确定文件名在<a>标签文本里,可直接定位链接再回溯到section:

from bs4 import BeautifulSoup
import re

target_link = soup.find('a', href=True, text=re.compile(r'CIcyano'))
if target_link:
    target_section = target_link.find_parent('section', class_='onecol habonecol')
    download_url = target_link['href']
    print(download_url)

内容的提问来源于stack exchange,提问作者Koelker12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 15:05:17