You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取HTML列表中所有层级的<li>标签

需求:按H1标题归类提取嵌套列表中的所有LI标签

我有一份带层级列表的文档,转成HTML后存在嵌套的<ul>结构。因为列表是嵌套的,没法提取外层和内层<ul>里的所有<li>标签,想要把每个<h1>标题对应的所有<li>标签归类到对应列表里。试过find、find_all、find_all_next这些方法都没成功,求解决方案。

原文档示例

HEADER1

  • The virus killed 56 people
  • Global press highlights hundreds of dogs jumping
    • A Twitter user posts photos of cats
HEADER2
  • President Biden talks on Sunday.

  • Hello World

  • How can I help you?

Header 3
  • The war in Gaza continues
  • Global press highlights best pizza
    • A Twitter user posts sushi
    • A Twitter user posts candy

转换后的HTML示例

<html>
 <body>
 <h1>HEADER1</h1>
  <ul>
   <li>
    the virus killed 56
   </li>
   <li>
    Global press
    <a href="https://www.example.com">
     highlight
    </a>
    hundreds of dogs jumping
    <ul>
     <li>
      A Twitter user
      <a href="http://example.com/xad/status/sda">
       posts
      </a>
      photos of cats
     </li>
    </ul>
   </li>
  </ul>
 <h1>HEADER2</h1>
  <ul>
   <li>
    President Biden talks on Sunday.
   </li>
   Hello World
   <li>
    How can I help you?
   </li>
  </ul>
 <h1>HEADER3</h1>
 <ul>
   <li>
    The war in Gaza continues
   </li>
   <li>
    Global press highlights best pizza
    <ul>
     <li>
      A Twitter user posts sushi
     </li>
     <li>
      A Twitter user posts candy
     </li>
    </ul>
   </li>
  </ul>
 </body>
</html>

预期结果示例

header1_posts = [<li>the virus killed 56</li>, <li>Global press<a href="https://www.example.com">highlight</a>hundreds of dogs jumping</li>, <li>A Twitter user<a href="http://example.com/xad/status/sda">posts</a>photos of cats</li>]
header2_posts = [...]
header3_posts = [...]

解决方案

用BeautifulSoup的递归查找能力提取所有层级的<li>标签,同时按<h1>分组处理:

代码实现

from bs4 import BeautifulSoup

html_content = """
<html>
 <body>
 <h1>HEADER1</h1>
  <ul>
   <li>
    the virus killed 56
   </li>
   <li>
    Global press
    <a href="https://www.example.com">
     highlight
    </a>
    hundreds of dogs jumping
    <ul>
     <li>
      A Twitter user
      <a href="http://example.com/xad/status/sda">
       posts
      </a>
      photos of cats
     </li>
    </ul>
   </li>
  </ul>
 <h1>HEADER2</h1>
  <ul>
   <li>
    President Biden talks on Sunday.
   </li>
   Hello World
   <li>
    How can I help you?
   </li>
  </ul>
 <h1>HEADER3</h1>
 <ul>
   <li>
    The war in Gaza continues
   </li>
   <li>
    Global press highlights best pizza
    <ul>
     <li>
      A Twitter user posts sushi
     </li>
     <li>
      A Twitter user posts candy
     </li>
    </ul>
   </li>
  </ul>
 </body>
</html>
"""

soup = BeautifulSoup(html_content, 'html.parser')
result = {}

# 遍历所有h1标签,按标题分组
for h1 in soup.find_all('h1'):
    header_key = h1.get_text(strip=True).lower() + '_posts'
    current_content = []
    
    # 获取当前h1之后的元素,直到遇到下一个h1
    for sibling in h1.find_next_siblings():
        if sibling.name == 'h1':
            break
        current_content.append(sibling)
    
    # 递归提取所有层级的li标签
    all_li = []
    for elem in current_content:
        all_li.extend(elem.find_all('li', recursive=True))
    
    # 清理li的HTML格式,去除多余空白
    cleaned_li = []
    for li in all_li:
        cleaned_str = str(li).replace('\n', '').replace('\t', '').strip()
        cleaned_li.append(cleaned_str)
    
    result[header_key] = cleaned_li

# 输出结果
for key, value in result.items():
    print(f"{key} = {value}")

代码说明

  1. 以每个<h1>为分组节点,获取其后续所有兄弟元素,直到下一个<h1>出现,确保只处理当前标题对应的内容块。
  2. 用find_all('li', recursive=True)递归提取内容块中所有层级的<li>标签,不管嵌套深度。
  3. 清理<li>的HTML字符串,去除多余换行和制表符,保留完整标签结构。
  4. 最终结果按headerX_posts的格式存储,符合预期需求。

内容的提问来源于stack exchange,提问作者Yoni Levine

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 04:15:21