如何从Python DataFrame列提取纯文本去除HTML标签并解决报错
问题原因
你的free text字段中存在空值(Pandas 中的 NaN / Python 中的 None),当 apply 遍历到这些空值时,BeautifulSoup 接收的输入参数为空,解析过程触发类型错误。
解决方案
根据使用场景可选以下方案:
- 方案1:兼容空值,保留原有逻辑,提取所有文本段为列表
from bs4 import BeautifulSoup import pandas as pd df['clean_text_list'] = df['free text'].apply( lambda x: list(BeautifulSoup(x, "html.parser").stripped_strings) if not pd.isna(x) else [] )
- 方案2:兼容空值,直接输出拼接完成的完整纯文本(更常用的场景)
from bs4 import BeautifulSoup import pandas as pd # separator参数指定不同标签文本之间的分隔符,strip参数自动去除首尾空白 df['clean_text'] = df['free text'].apply( lambda x: BeautifulSoup(x, "html.parser").get_text(strip=True, separator=' ') if not pd.isna(x) else '' )
- 可选前置操作:如果不需要保留空值对应的行,可以先删除空值后直接运行原有代码
# 先删除free text列为空的行 df = df.dropna(subset=['free text']) # 之后可直接运行你原来的代码 df['clean_text_list'] = df['free text'].apply( lambda x: list(BeautifulSoup(x, "html.parser").stripped_strings) )
内容的提问来源于stack exchange,提问作者adey27
相关产品推荐
相关产品推荐

