如何用BeautifulSoup从HTML中精准提取Part1对应的列表内容?
精准提取Part1对应的列表项内容
需求
从以下HTML代码中,提取**Part1:**对应的列表项文本(即"This is part one."和"Please extract me"):
<div style="margin-bottom:2px;"><strong>Part1:</strong></div> <ul> <li>This is part one.</li> <li>Please extract me</li> </ul> <div style="margin-bottom:2px;"><strong>PartTwo:</strong></div> <ul> <li>This is part 2</li> <li>This has not to be extracted</li> </ul>
之前的问题
- 仅匹配文本的代码只能拿到"Part1:"本身,无法关联后续列表:
soup = BeautifulSoup(html_document, 'html.parser') part1 = soup(text=lambda t: "Part1:" in t.text) - 遍历所有
<li>的代码会同时包含PartTwo的内容,无法精准筛选:for ul in soup: for li in soup.findAll('li'): print(li)
正确实现代码
from bs4 import BeautifulSoup html_document = """ <div style="margin-bottom:2px;"><strong>Part1:</strong></div> <ul> <li>This is part one.</li> <li>Please extract me</li> </ul> <div style="margin-bottom:2px;"><strong>PartTwo:</strong></div> <ul> <li>This is part 2</li> <li>This has not to be extracted</li> </ul> """ soup = BeautifulSoup(html_document, 'html.parser') # 定位到文本为"Part1:"的strong标签 part1_strong = soup.find('strong', string="Part1:") # 获取该标签父div的下一个兄弟ul节点 part1_list = part1_strong.parent.find_next_sibling('ul') # 提取所有li的文本并清理空格 part1_items = [li.get_text(strip=True) for li in part1_list.find_all('li')] print(part1_items) # 输出结果: ['This is part one.', 'Please extract me']
代码说明
- 精准定位标签:用
soup.find('strong', string="Part1:")直接找到目标<strong>,避免模糊匹配带来的干扰 - 关联对应列表:通过
parent找到<strong>所在的<div>,再用find_next_sibling('ul')获取紧邻的后续列表,确保只取Part1对应的内容 - 提取清理文本:遍历列表内的
<li>,用get_text(strip=True)提取文本并自动去除前后多余空格
内容的提问来源于stack exchange,提问作者question12
相关产品推荐
相关产品推荐

