Python中使用BeautifulSoup(BS4)仅提取单个标签自身文本内容的解决方案
解决BeautifulSoup提取单个标签自身文本(排除子标签内容)的问题
嘿,我之前也遇到过这个问题!get_text()默认会递归提取目标标签下所有子节点的文本,包括嵌套在里面的子标签内容,所以才会把那些你不需要的文本也带出来。这里有两种靠谱的解决方法:
方法1:遍历直接子节点,筛选文本类型节点
通过contents获取目标标签的所有直接子节点,只保留字符串类型的节点,再处理掉多余空格和引号后拼接:
from bs4 import BeautifulSoup html = '''<div class="test__title" data-intro="intro"> "The first text I need" "The second text I need" <p class="test__title__promt"> "I do not need it" "I do not need it" <p class="test__title__second__promt"> "I do not need it" "I do not need it"</div>''' soup = BeautifulSoup(html, 'html.parser') test_title = soup.find('div', class_='test__title') # 筛选直接子节点中的文本,去掉无效空格和引号后拼接 target_text = ' '.join([node.strip().strip('"') for node in test_title.contents if isinstance(node, str)]) print(target_text)
输出结果:The first text I need The second text I need
方法2:使用find_all的recursive=False参数
直接指定recursive=False,只提取当前标签的直接文本节点,不深入子标签:
from bs4 import BeautifulSoup html = '''<div class="test__title" data-intro="intro"> "The first text I need" "The second text I need" <p class="test__title__promt"> "I do not need it" "I do not need it" <p class="test__title__second__promt"> "I do not need it" "I do not need it"</div>''' soup = BeautifulSoup(html, 'html.parser') test_title = soup.find('div', class_='test__title') # 只获取当前标签的直接文本,不递归子标签 target_text = ' '.join([s.strip().strip('"') for s in test_title.find_all(text=True, recursive=False)]) print(target_text)
输出结果同样符合你的需求。
补充说明
原来的get_text()方法默认recursive=True,会遍历所有嵌套的子标签并提取文本,所以才会包含那些你不需要的内容。上面两种方法都是通过限制提取范围,只获取目标标签自身的直接文本节点,完美解决你的问题~
内容的提问来源于stack exchange,提问作者mooncdx
相关产品推荐
相关产品推荐

