You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的BeautifulSoup提取嵌套<li>标签中的无重复文本

解决嵌套<li>重复提取与缩进保留问题

问题分析

使用find_all('li')会抓取所有层级的<li>标签,包括嵌套在父<li>里的子<li>,导致内容重复;直接取.text会保留多余空格且无法体现层级缩进。

解决方案

通过递归遍历层级处理<li>元素,按层级添加缩进,同时用stripped_strings清理多余空格,确保内容仅提取一次且格式清晰。

完整代码

#! python3
import requests
from bs4 import BeautifulSoup

url = 'https://numismatics.org/pella/id/price.1?lang=it'
html = requests.get(url)
s = BeautifulSoup(html.content, 'html.parser')

# 定位目标区域
description = s.find(class_='metadata_section')

# 存储结果的列表
output_content = []

def process_list_item(li_element, indent_level=0):
    # 提取并清理当前li的文本(去除多余空格、换行)
    cleaned_text = ' '.join(li_element.stripped_strings)
    if cleaned_text:
        # 根据层级添加缩进,用4个空格代表一级
        indented_text = '    ' * indent_level + cleaned_text
        output_content.append(indented_text)
    
    # 递归处理当前li下的直接子li(避免重复抓取深层嵌套)
    for child_li in li_element.find_all('li', recursive=False):
        process_list_item(child_li, indent_level + 1)

# 处理目标区域下的所有顶级li
if description:
    top_level_lis = description.find_all('li', recursive=False)
    for li in top_level_lis:
        process_list_item(li)

# 打印结果
for line in output_content:
    print(line)

# 保存到txt文件
with open('coin_description.txt', 'w', encoding='utf-8') as file:
    file.write('\n'.join(output_content))

代码说明

  1. 递归遍历:process_list_item函数处理单个<li>,再递归处理其直接子<li>,确保每个元素仅被处理一次
  2. 缩进控制:通过indent_level参数控制缩进,每深入一层增加4个空格
  3. 文本清理:stripped_strings自动去除文本中的多余空格、换行和制表符,用空格连接成整洁的文本
  4. 结果保存:将处理后的内容存入列表,既可以打印也能写入文件

内容的提问来源于stack exchange,提问作者user3610033

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 22:33:22