You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup从HTML中精准提取Part1对应的列表内容?

精准提取Part1对应的列表项内容

需求

从以下HTML代码中,提取**Part1:**对应的列表项文本(即"This is part one."和"Please extract me"):

<div style="margin-bottom:2px;"><strong>Part1:</strong></div>
<ul>
<li>This is part one.</li>
<li>Please extract me</li>
</ul>
&nbsp;
<div style="margin-bottom:2px;"><strong>PartTwo:</strong></div>
<ul>
<li>This is part 2</li>
<li>This has not to be extracted</li>

</ul>

之前的问题

  • 仅匹配文本的代码只能拿到"Part1:"本身,无法关联后续列表:
    soup = BeautifulSoup(html_document, 'html.parser')
    part1 = soup(text=lambda t: "Part1:" in t.text)
    
  • 遍历所有<li>的代码会同时包含PartTwo的内容,无法精准筛选:
    for ul in soup:
        for li in soup.findAll('li'):
            print(li)
    

正确实现代码

from bs4 import BeautifulSoup

html_document = """
<div style="margin-bottom:2px;"><strong>Part1:</strong></div>
<ul>
<li>This is part one.</li>
<li>Please extract me</li>
</ul>
&nbsp;
<div style="margin-bottom:2px;"><strong>PartTwo:</strong></div>
<ul>
<li>This is part 2</li>
<li>This has not to be extracted</li>

</ul>
"""

soup = BeautifulSoup(html_document, 'html.parser')

# 定位到文本为"Part1:"的strong标签
part1_strong = soup.find('strong', string="Part1:")
# 获取该标签父div的下一个兄弟ul节点
part1_list = part1_strong.parent.find_next_sibling('ul')
# 提取所有li的文本并清理空格
part1_items = [li.get_text(strip=True) for li in part1_list.find_all('li')]

print(part1_items)
# 输出结果: ['This is part one.', 'Please extract me']

代码说明

  1. 精准定位标签:用soup.find('strong', string="Part1:")直接找到目标<strong>,避免模糊匹配带来的干扰
  2. 关联对应列表:通过parent找到<strong>所在的<div>,再用find_next_sibling('ul')获取紧邻的后续列表,确保只取Part1对应的内容
  3. 提取清理文本:遍历列表内的<li>,用get_text(strip=True)提取文本并自动去除前后多余空格

内容的提问来源于stack exchange,提问作者question12

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 11:03:38