You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除DataFrame指定列中的<p><br>等HTML转义标签仅保留正文

解决方案

你需要分两步处理:先还原HTML转义字符,再移除所有HTML标签保留正文。可以用Python标准库html做转义还原,搭配BeautifulSoup做HTML标签提取,是兼容性最好的方案。

依赖安装

如果还没有安装对应的依赖,执行:

pip install pandas beautifulsoup4

实现代码

import pandas as pd
import html
from bs4 import BeautifulSoup

def clean_html_content(raw_str):
    # 还原所有HTML转义字符
    unescaped_str = html.unescape(raw_str)
    # 解析HTML提取纯文本
    soup = BeautifulSoup(unescaped_str, "html.parser")
    # 返回提取到的文本,strip=True可同时移除首尾多余空白
    return soup.get_text()

# 对Description列应用清理函数
df["Description"] = df["Description"].apply(clean_html_content)

轻量替代方案(无需安装BeautifulSoup)

如果不想引入额外依赖,也可以用正则表达式完成标签清理,适合简单场景:

import pandas as pd
import html
import re

def clean_html_content(raw_str):
    unescaped_str = html.unescape(raw_str)
    # 正则匹配所有HTML标签并移除
    return re.sub(r"<[^>]*>", "", unescaped_str).strip()

df["Description"] = df["Description"].apply(clean_html_content)

以上两种方案都可以直接匹配你给出的样例输入,输出完全符合预期。

内容的提问来源于stack exchange,提问作者aditya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 22:15:04