如何使用Beautiful Soup提取深度嵌套的<p>标签文本内容
解决方案
使用BeautifulSoup实现
核心思路是先定位到所有目标p标签共同的顶层父容器top-panel,再递归查找该容器下所有层级的p标签,不受中间嵌套结构限制,代码示例如下:
from bs4 import BeautifulSoup # 替换为你的实际HTML内容 html_content = """ <div class="top-panel"> <div class="inside-panel-0"> <h1 class="h1-title">Some Title</h1> </div> <div class="inside-panel-0"> <div class="inside-panel-1"> <p> I want to extract this copy</p> </div> <div class="inside-panel-1"> <p>I want to extract this copy</p> </div> </div> </div> """ soup = BeautifulSoup(html_content, 'html.parser') # 定位顶层容器,查找其下所有p标签 target_p = soup.find('div', class_='top-panel').find_all('p') # 也可以用css选择器写法:target_p = soup.select('div.top-panel p') # 提取文本 result = [p.get_text(strip=True) for p in target_p] print(result)
运行后输出结果为:
['I want to extract this copy', 'I want to extract this copy']
其他可选方案:lxml+xpath
如果你追求更高的解析效率,也可以用lxml库的xpath语法实现跨层级查找,写法更简洁:
from lxml import etree html = etree.HTML(html_content) # 匹配top-panel下所有层级的p标签文本 result = [text.strip() for text in html.xpath('//div[@class="top-panel"]//p/text()') if text.strip()] print(result)
你之前的方案失效是因为仅匹配了固定层级的直接子p标签,上述两种方法均采用全层级递归匹配逻辑,不受p标签上层嵌套结构的影响,只要在指定的顶层容器范围内即可被识别。
内容的提问来源于stack exchange,提问作者Freddy
相关产品推荐
相关产品推荐

