如何在Beautiful Soup的lambda筛选中匹配含Investment的h3标签并提取金额
问题:匹配包含指定文本的h3标签并提取金额
目标HTML内容
html_content2 =""" <h3 style="cear: both;"> <abbr title="European Union">EU</abbr>Investment</h3> <div class="conditions"> <p>bla bla bla </p> </div> <p style="margin-bottom: 0;"> <span class="amount">66000 €</span> </p>"""
现有代码
from bs4 import BeautifulSoup html_content=html_content1 soup = BeautifulSoup(html_content, "lxml") t3 = soup.find(lambda tag:tag.name=="h3" and ": Investment").find_next_sibling().find_next_sibling("p").find("span").contents print(t3)
问题分析
你需要匹配包含"Investment"文本的h3标签,但之前的tag.contents==": Investment"写法无法生效,原因是:
tag.contents返回的是h3内部的子节点列表(包含<abbr>标签对象和字符串"Investment"),无法直接与字符串比较- 目标h3的纯文本是"EUInvestment",并不包含冒号,匹配文本存在错误
解决方法
方法1:修正lambda匹配逻辑
使用tag.get_text(strip=True)获取h3的纯文本内容,再判断是否包含"Investment",同时优化节点定位逻辑:
from bs4 import BeautifulSoup html_content = html_content2 # 替换为正确的变量名html_content2 soup = BeautifulSoup(html_content, "lxml") # 匹配包含Investment文本的h3标签 target_h3 = soup.find(lambda tag: tag.name == "h3" and "Investment" in tag.get_text(strip=True)) if target_h3: # 定位到h3之后的div节点,再找其下一个兄弟节点p target_p = target_h3.find_next_sibling("div").find_next_sibling("p") if target_p: amount_span = target_p.find("span", class_="amount") if amount_span: print(amount_span.get_text(strip=True)) # 输出:66000 €
方法2:使用CSS选择器简化代码
利用BeautifulSoup的CSS选择器直接定位,代码更简洁:
from bs4 import BeautifulSoup html_content = html_content2 soup = BeautifulSoup(html_content, "lxml") # 选择器逻辑:找到包含Investment的h3,然后取其后的div的下一个p标签内的.amount元素 amount = soup.select_one('h3:contains("Investment") ~ div ~ p .amount').get_text(strip=True) print(amount) # 输出:66000 €
关键说明
get_text(strip=True):提取标签的纯文本内容并去除首尾空白,适合处理包含子标签的文本匹配场景- CSS选择器中的
~:表示选择当前元素之后的同级兄弟元素 :contains("文本"):匹配包含指定文本的标签(需配合lxml解析器使用)
内容的提问来源于stack exchange,提问作者JFerro
相关产品推荐
相关产品推荐

