如何用BeautifulSoup提取HTML中content与后续标签间的文本?
提取content与后续同类标题块之间的文本(Python实现)
用Python的BeautifulSoup库就能搞定这个需求,步骤和代码如下:
首先安装依赖:
pip install beautifulsoup4
核心思路
- 解析HTML内容,定位到包含
content:的<strong>标签所在的<p>块 - 从这个
<p>块开始,遍历后续所有同级<p>元素 - 收集每个
<p>的文本,直到遇到下一个包含带冒号的标题(比如time:、location:这类)的<p><strong>块为止
示例代码
from bs4 import BeautifulSoup # 替换成你的实际HTML内容 html = """ <p><strong> name:</strong></p> <p> sdfsdf </p> <p><strong> content:</strong></p> <p> yangben</p> <p> dsfs </p> <p> dfsds </p> <p> sdfs </p> <p><strong> time:</strong></p> <p> 2020-10-10</p> <p><strong> ll:</strong></p> <p> 2020-10-10</p> """ soup = BeautifulSoup(html, 'html.parser') # 定位content标题块 content_header = soup.find('strong', string=lambda t: t and 'content:' in t.strip()) if not content_header: print("找不到content对应的标题块") else: current_p = content_header.parent result = [] # 遍历后续兄弟p元素 for sibling in current_p.find_next_siblings('p'): # 检查是否是下一个标题块(带冒号的strong) strong = sibling.find('strong') if strong and ':' in strong.text.strip(): break # 提取非空文本 text = sibling.text.strip() if text: result.append(text) # 输出结果 print("提取到的内容:") for line in result: print(line)
代码说明
- 用
lambda表达式匹配content:,能忽略文本前后的空格,避免匹配失败 - 判断停止条件时,只要遇到带冒号的
<strong>标题就停止,兼容time、location等任何同类标签 - 自动过滤空文本,只保留有效内容
运行结果
提取到的内容: yangben dsfs dfsds sdfs
内容的提问来源于stack exchange,提问作者Lei Hao
相关产品推荐
相关产品推荐

