You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中使用BeautifulSoup(BS4)仅提取单个标签自身文本内容的解决方案

解决BeautifulSoup提取单个标签自身文本(排除子标签内容)的问题

嘿,我之前也遇到过这个问题!get_text()默认会递归提取目标标签下所有子节点的文本,包括嵌套在里面的子标签内容,所以才会把那些你不需要的文本也带出来。这里有两种靠谱的解决方法:

方法1:遍历直接子节点,筛选文本类型节点

通过contents获取目标标签的所有直接子节点,只保留字符串类型的节点,再处理掉多余空格和引号后拼接:

from bs4 import BeautifulSoup

html = '''<div class="test__title" data-intro="intro"> "The first text I need" "The second text I need" <p class="test__title__promt"> "I do not need it" "I do not need it" <p class="test__title__second__promt"> "I do not need it" "I do not need it"</div>'''
soup = BeautifulSoup(html, 'html.parser')

test_title = soup.find('div', class_='test__title')
# 筛选直接子节点中的文本,去掉无效空格和引号后拼接
target_text = ' '.join([node.strip().strip('"') for node in test_title.contents if isinstance(node, str)])
print(target_text)

输出结果:The first text I need The second text I need

方法2:使用find_all的recursive=False参数

直接指定recursive=False,只提取当前标签的直接文本节点,不深入子标签:

from bs4 import BeautifulSoup

html = '''<div class="test__title" data-intro="intro"> "The first text I need" "The second text I need" <p class="test__title__promt"> "I do not need it" "I do not need it" <p class="test__title__second__promt"> "I do not need it" "I do not need it"</div>'''
soup = BeautifulSoup(html, 'html.parser')

test_title = soup.find('div', class_='test__title')
# 只获取当前标签的直接文本,不递归子标签
target_text = ' '.join([s.strip().strip('"') for s in test_title.find_all(text=True, recursive=False)])
print(target_text)

输出结果同样符合你的需求。

补充说明

原来的get_text()方法默认recursive=True,会遍历所有嵌套的子标签并提取文本,所以才会包含那些你不需要的内容。上面两种方法都是通过限制提取范围,只获取目标标签自身的直接文本节点,完美解决你的问题~

内容的提问来源于stack exchange,提问作者mooncdx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 02:49:05