You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取HTML字符串中的各部分内容并进行分组?

嘿,我来帮你搞定从这段HTML里提取内容并分组的事儿!首先咱们先把你提供的HTML片段整理完整(补全了截断部分方便演示),再一步步拆解提取逻辑:

从HTML片段提取分组内容的实用方案

先看完整的目标HTML片段

<div xmlns="http://www.w3.org/1999/xhtml" class="filters-container"> 
  <div class="search-filter" id="search-filter-type">
    <span class="filter-option-title">Car Type</span> 
    <ul> 
      <li> 
        <input id="filter-types-all" name="filter-types-all" type="checkbox" checked="checked" /> 
        <label for="filter-types-all"><span class="filter-label">All Types</span></label>
      </li>
      <!-- 假设还有其他车型选项 -->
      <li> 
        <input id="filter-types-sedan" name="filter-types" type="checkbox" /> 
        <label for="filter-types-sedan"><span class="filter-label">Sedan</span></label>
      </li>
      <li> 
        <input id="filter-types-suv" name="filter-types" type="checkbox" /> 
        <label for="filter-types-suv"><span class="filter-label">SUV</span></label>
      </li>
    </ul>
  </div>
</div>

我会用Python的BeautifulSoup库来做解析,这是处理HTML提取最常用的工具之一,逻辑清晰易上手。

1. 先安装依赖

如果还没装的话,先执行这条命令:

pip install beautifulsoup4

2. 解析+分组提取的代码示例

这段代码会把内容按「过滤器标题」和「选项组」自动分组,输出结构化的结果:

from bs4 import BeautifulSoup

# 把你的HTML字符串放这里
html_content = """
<div xmlns="http://www.w3.org/1999/xhtml" class="filters-container"> 
  <div class="search-filter" id="search-filter-type">
    <span class="filter-option-title">Car Type</span> 
    <ul> 
      <li> 
        <input id="filter-types-all" name="filter-types-all" type="checkbox" checked="checked" /> 
        <label for="filter-types-all"><span class="filter-label">All Types</span></label>
      </li>
      <li> 
        <input id="filter-types-sedan" name="filter-types" type="checkbox" /> 
        <label for="filter-types-sedan"><span class="filter-label">Sedan</span></label>
      </li>
      <li> 
        <input id="filter-types-suv" name="filter-types" type="checkbox" /> 
        <label for="filter-types-suv"><span class="filter-label">SUV</span></label>
      </li>
    </ul>
  </div>
</div>
"""

# 初始化解析器
soup = BeautifulSoup(html_content, "html.parser")

# 用来存储分组后的结果
filtered_groups = []

# 遍历每个过滤器模块
for filter_block in soup.find_all("div", class_="search-filter"):
    # 提取过滤器标题
    filter_title = filter_block.find("span", class_="filter-option-title").text.strip()
    # 提取所有选项
    option_list = []
    for option_item in filter_block.find("ul").find_all("li"):
        input_tag = option_item.find("input")
        option_label = option_item.find("span", class_="filter-label").text.strip()
        option_list.append({
            "选项名称": option_label,
            "输入框ID": input_tag.get("id"),
            "是否选中": input_tag.get("checked") is not None
        })
    # 把标题和选项组打包
    filtered_groups.append({
        "过滤器标题": filter_title,
        "选项集合": option_list
    })

# 打印最终分组结果
for group in filtered_groups:
    print(f"📌 {group['过滤器标题']}")
    print("---")
    for opt in group["选项集合"]:
        status = "✅ 已选中" if opt["是否选中"] else "🔲 未选中"
        print(f"- {opt['选项名称']} {status}")

3. 运行后的输出效果

你会得到清晰的分组结果:

📌 Car Type

  • All Types ✅ 已选中
  • Sedan 🔲 未选中
  • SUV 🔲 未选中

如果是用JavaScript的话,也可以用DOMParser或者Cheerio库实现类似逻辑,核心思路都是通过元素的class、id或标签定位,再提取对应的文本和属性。

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:49:00