You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Python DataFrame列提取纯文本去除HTML标签并解决报错

问题原因

你的free text字段中存在空值(Pandas 中的 NaN / Python 中的 None),当 apply 遍历到这些空值时,BeautifulSoup 接收的输入参数为空,解析过程触发类型错误。

解决方案

根据使用场景可选以下方案:

  • 方案1:兼容空值,保留原有逻辑,提取所有文本段为列表
from bs4 import BeautifulSoup
import pandas as pd

df['clean_text_list'] = df['free text'].apply(
    lambda x: list(BeautifulSoup(x, "html.parser").stripped_strings) if not pd.isna(x) else []
)
  • 方案2:兼容空值,直接输出拼接完成的完整纯文本(更常用的场景)
from bs4 import BeautifulSoup
import pandas as pd

# separator参数指定不同标签文本之间的分隔符,strip参数自动去除首尾空白
df['clean_text'] = df['free text'].apply(
    lambda x: BeautifulSoup(x, "html.parser").get_text(strip=True, separator=' ') if not pd.isna(x) else ''
)
  • 可选前置操作:如果不需要保留空值对应的行,可以先删除空值后直接运行原有代码
# 先删除free text列为空的行
df = df.dropna(subset=['free text'])
# 之后可直接运行你原来的代码
df['clean_text_list'] = df['free text'].apply(
    lambda x: list(BeautifulSoup(x, "html.parser").stripped_strings)
)

内容的提问来源于stack exchange,提问作者adey27

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 19:54:01