You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup查找包含指定文本及子元素的<li>标签?

解决方法

你之前的方法失效,是因为当<li>包含<a>这类子元素时,该<li>的string属性会返回None,导致基于string参数的匹配无法生效。以下是几种可行的查询方式:

方法一:Lambda表达式同时检查文本与子元素

直接通过lambda表达式,既验证<li>的文本内容包含目标字符串,又确认它存在<a>子元素:

from bs4 import BeautifulSoup

html = '''
<li>
   Country:
   <a href="example.com">Germany</a>
</li>
'''
soup = BeautifulSoup(html, 'html.parser')

target_li = soup.find("li", lambda tag: "Country: " in tag.get_text() and tag.find("a") is not None)

如果需要忽略文本中的空格和换行,可以把tag.get_text()改成tag.get_text(strip=True),只要目标字符串的核心内容存在即可。

方法二:先匹配文本再筛选子元素

先找到所有包含目标文本的<li>,再从中过滤出带有<a>子元素的标签:

import re
from bs4 import BeautifulSoup

soup = BeautifulSoup(html, 'html.parser')
# 匹配包含"Country: "文本的<li>(正则会检查所有子文本节点)
candidates = soup.find_all("li", string=re.compile(r'Country: '))
# 筛选出有<a>子元素的结果
target_li = next((li for li in candidates if li.find("a")), None)

方法三:CSS选择器结合子元素检查

先用CSS选择器定位包含目标文本的<li>,再验证它是否存在<a>子元素:

from bs4 import BeautifulSoup

soup = BeautifulSoup(html, 'html.parser')
candidate_li = soup.select_one('li:contains("Country: ")')
target_li = candidate_li if candidate_li and candidate_li.find("a") else None

内容的提问来源于stack exchange,提问作者kutas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.15 16:39:59