You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python DataFrame将带HTML标签的技术数据拆分为多列

解决方案

你要做的是解析列中存储的HTML列表内容,将每个<li>标签内的键值对转换为单独的DataFrame列,推荐用HTML解析方案实现,比字符串硬拆分稳定性更高。

依赖安装(如未安装相关库)

pip install pandas beautifulsoup4

代码实现(BeautifulSoup 方案,推荐)

import pandas as pd
from bs4 import BeautifulSoup
import html

# 定义单条HTML内容的解析函数
def parse_tech_spec(html_str):
    # 解码HTML实体(把&lt;这类转义字符还原为<)
    decoded_html = html.unescape(html_str)
    # 解析HTML结构
    soup = BeautifulSoup(decoded_html, "html.parser")
    spec_dict = {}
    # 遍历所有li标签提取键值对
    for li_tag in soup.find_all("li"):
        text_content = li_tag.get_text(strip=True)
        # 按第一个冒号拆分属性名和属性值
        if ":" in text_content:
            attr_name, attr_value = text_content.split(":", 1)
            spec_dict[attr_name.strip()] = attr_value.strip()
    return pd.Series(spec_dict)

# 对目标列执行解析,合并到原DataFrame
parsed_specs = df["technische_daten"].apply(parse_tech_spec)
result_df = pd.concat([df, parsed_specs], axis=1)

纯正则实现(无需额外安装BeautifulSoup)

如果不想引入第三方HTML解析库,可以用正则匹配实现相同效果:

import pandas as pd
import html
import re

def parse_tech_spec_regex(html_str):
    decoded_html = html.unescape(html_str)
    # 匹配所有li标签内的内容
    li_items = re.findall(r"<li>(.*?)</li>", decoded_html, flags=re.DOTALL)
    spec_dict = {}
    for item in li_items:
        if ":" in item:
            attr_name, attr_value = item.split(":", 1)
            spec_dict[attr_name.strip()] = attr_value.strip()
    return pd.Series(spec_dict)

parsed_specs = df["technische_daten"].apply(parse_tech_spec_regex)
result_df = pd.concat([df, parsed_specs], axis=1)

说明

  • 最终result_df中会新增所有解析出来的属性列,缺失属性的行会自动填充NaN
  • 相比你之前用固定位置拆分的方法,该方案可以适配不同行数、不同属性顺序的HTML内容

内容的提问来源于stack exchange,提问作者akshay KATYAL

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 13:15:05