如何解决pandas DataFrame的HTML列用BeautifulSoup转纯文本报错问题
错误原因
BeautifulSoup仅支持传入单个HTML字符串进行解析,你直接传入整列的Series对象df['Body'],触发了pandas的Series真值校验规则,导致报错。
解决方案
方案1:使用BeautifulSoup逐行处理
适合数据量较小的场景,代码如下:
from bs4 import BeautifulSoup import pandas as pd def html_to_text(html_str): # 兼容空值场景,避免空值触发报错 if pd.isna(html_str): return "" soup = BeautifulSoup(html_str, "html.parser") # 若需要保留文本原有空格、换行,去掉strip=True参数即可 return soup.get_text(strip=True) # 转换结果存储到新列Body_纯文本 df["Body_纯文本"] = df["Body"].apply(html_to_text)
方案2:使用selectolax逐行处理
你代码中已经导入了selectolax库,该库解析性能远高于BeautifulSoup,适合数据量较大的场景,代码如下:
from selectolax.parser import HTMLParser import pandas as pd def html_to_text(html_str): if pd.isna(html_str): return "" tree = HTMLParser(html_str) # 提前移除无文本内容的无效标签,提升解析效率 tree.strip_tags(["script", "style", "meta", "link", "noscript"]) return tree.text(strip=True) df["Body_纯文本"] = df["Body"].apply(html_to_text)
处理完成后可以通过print(df["Body_纯文本"])查看转换后的结果。
内容的提问来源于stack exchange,提问作者edublog
相关产品推荐
相关产品推荐

