You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决pandas DataFrame的HTML列用BeautifulSoup转纯文本报错问题

错误原因

BeautifulSoup仅支持传入单个HTML字符串进行解析,你直接传入整列的Series对象df['Body'],触发了pandas的Series真值校验规则,导致报错。

解决方案

方案1:使用BeautifulSoup逐行处理

适合数据量较小的场景,代码如下:

from bs4 import BeautifulSoup
import pandas as pd

def html_to_text(html_str):
    # 兼容空值场景,避免空值触发报错
    if pd.isna(html_str):
        return ""
    soup = BeautifulSoup(html_str, "html.parser")
    # 若需要保留文本原有空格、换行,去掉strip=True参数即可
    return soup.get_text(strip=True)

# 转换结果存储到新列Body_纯文本
df["Body_纯文本"] = df["Body"].apply(html_to_text)

方案2:使用selectolax逐行处理

你代码中已经导入了selectolax库,该库解析性能远高于BeautifulSoup,适合数据量较大的场景,代码如下:

from selectolax.parser import HTMLParser
import pandas as pd

def html_to_text(html_str):
    if pd.isna(html_str):
        return ""
    tree = HTMLParser(html_str)
    # 提前移除无文本内容的无效标签,提升解析效率
    tree.strip_tags(["script", "style", "meta", "link", "noscript"])
    return tree.text(strip=True)

df["Body_纯文本"] = df["Body"].apply(html_to_text)

处理完成后可以通过print(df["Body_纯文本"])查看转换后的结果。

内容的提问来源于stack exchange,提问作者edublog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 12:24:07