如何用BeautifulSoup提取HTML列表中所有层级的<li>标签
需求:按H1标题归类提取嵌套列表中的所有LI标签
我有一份带层级列表的文档,转成HTML后存在嵌套的<ul>结构。因为列表是嵌套的,没法提取外层和内层<ul>里的所有<li>标签,想要把每个<h1>标题对应的所有<li>标签归类到对应列表里。试过find、find_all、find_all_next这些方法都没成功,求解决方案。
原文档示例
HEADER1
HEADER2
- The virus killed 56 people
- Global press highlights hundreds of dogs jumping
- A Twitter user posts photos of cats
Header 3
President Biden talks on Sunday.
Hello World
How can I help you?
- The war in Gaza continues
- Global press highlights best pizza
- A Twitter user posts sushi
- A Twitter user posts candy
转换后的HTML示例
<html> <body> <h1>HEADER1</h1> <ul> <li> the virus killed 56 </li> <li> Global press <a href="https://www.example.com"> highlight </a> hundreds of dogs jumping <ul> <li> A Twitter user <a href="http://example.com/xad/status/sda"> posts </a> photos of cats </li> </ul> </li> </ul> <h1>HEADER2</h1> <ul> <li> President Biden talks on Sunday. </li> Hello World <li> How can I help you? </li> </ul> <h1>HEADER3</h1> <ul> <li> The war in Gaza continues </li> <li> Global press highlights best pizza <ul> <li> A Twitter user posts sushi </li> <li> A Twitter user posts candy </li> </ul> </li> </ul> </body> </html>
预期结果示例
header1_posts = [<li>the virus killed 56</li>, <li>Global press<a href="https://www.example.com">highlight</a>hundreds of dogs jumping</li>, <li>A Twitter user<a href="http://example.com/xad/status/sda">posts</a>photos of cats</li>] header2_posts = [...] header3_posts = [...]
解决方案
用BeautifulSoup的递归查找能力提取所有层级的<li>标签,同时按<h1>分组处理:
代码实现
from bs4 import BeautifulSoup html_content = """ <html> <body> <h1>HEADER1</h1> <ul> <li> the virus killed 56 </li> <li> Global press <a href="https://www.example.com"> highlight </a> hundreds of dogs jumping <ul> <li> A Twitter user <a href="http://example.com/xad/status/sda"> posts </a> photos of cats </li> </ul> </li> </ul> <h1>HEADER2</h1> <ul> <li> President Biden talks on Sunday. </li> Hello World <li> How can I help you? </li> </ul> <h1>HEADER3</h1> <ul> <li> The war in Gaza continues </li> <li> Global press highlights best pizza <ul> <li> A Twitter user posts sushi </li> <li> A Twitter user posts candy </li> </ul> </li> </ul> </body> </html> """ soup = BeautifulSoup(html_content, 'html.parser') result = {} # 遍历所有h1标签,按标题分组 for h1 in soup.find_all('h1'): header_key = h1.get_text(strip=True).lower() + '_posts' current_content = [] # 获取当前h1之后的元素,直到遇到下一个h1 for sibling in h1.find_next_siblings(): if sibling.name == 'h1': break current_content.append(sibling) # 递归提取所有层级的li标签 all_li = [] for elem in current_content: all_li.extend(elem.find_all('li', recursive=True)) # 清理li的HTML格式,去除多余空白 cleaned_li = [] for li in all_li: cleaned_str = str(li).replace('\n', '').replace('\t', '').strip() cleaned_li.append(cleaned_str) result[header_key] = cleaned_li # 输出结果 for key, value in result.items(): print(f"{key} = {value}")
代码说明
- 以每个
<h1>为分组节点,获取其后续所有兄弟元素,直到下一个<h1>出现,确保只处理当前标题对应的内容块。 - 用
find_all('li', recursive=True)递归提取内容块中所有层级的<li>标签,不管嵌套深度。 - 清理
<li>的HTML字符串,去除多余换行和制表符,保留完整标签结构。 - 最终结果按
headerX_posts的格式存储,符合预期需求。
内容的提问来源于stack exchange,提问作者Yoni Levine
相关产品推荐
相关产品推荐

