如何移除DataFrame指定列中的<p><br>等HTML转义标签仅保留正文
解决方案
你需要分两步处理:先还原HTML转义字符,再移除所有HTML标签保留正文。可以用Python标准库html做转义还原,搭配BeautifulSoup做HTML标签提取,是兼容性最好的方案。
依赖安装
如果还没有安装对应的依赖,执行:
pip install pandas beautifulsoup4
实现代码
import pandas as pd import html from bs4 import BeautifulSoup def clean_html_content(raw_str): # 还原所有HTML转义字符 unescaped_str = html.unescape(raw_str) # 解析HTML提取纯文本 soup = BeautifulSoup(unescaped_str, "html.parser") # 返回提取到的文本,strip=True可同时移除首尾多余空白 return soup.get_text() # 对Description列应用清理函数 df["Description"] = df["Description"].apply(clean_html_content)
轻量替代方案(无需安装BeautifulSoup)
如果不想引入额外依赖,也可以用正则表达式完成标签清理,适合简单场景:
import pandas as pd import html import re def clean_html_content(raw_str): unescaped_str = html.unescape(raw_str) # 正则匹配所有HTML标签并移除 return re.sub(r"<[^>]*>", "", unescaped_str).strip() df["Description"] = df["Description"].apply(clean_html_content)
以上两种方案都可以直接匹配你给出的样例输入,输出完全符合预期。
内容的提问来源于stack exchange,提问作者aditya
相关产品推荐
相关产品推荐

