如何移除Pandas DataFrame列表中的特定字符与空白符?
问题描述
我通过以下代码读取CSV文件为Pandas DataFrame:
import pandas as pd from ast import literal_eval df = pd.read_csv('listing.csv')
删除所有无描述的行并重置索引:
no_description = df['Description'].isna() descriptions = df[~no_description].reset_index()
修正某行偏移的数据(左移一位):
descriptions.iloc[538, -17:] = descriptions.iloc[538, -17:].shift(-1)
Description列看起来是列表,但读取后是字符串格式,示例数据如下:
| Index | Category | Title | Description |
|---|---|---|---|
| 1 | Treehouse | Red Kite Tree Tent | ['About this Space', 'On the Domaine de l 'Arbre', 'in Cabane\n\n'] |
| 2 | Treehouse | Nature Cabin | ['About this Space', 'Nature is everywhere\n\nA', 'treehouse to indulge, relax and go where no one will be'] |
我用literal_eval将其转换为真正的列表:
df['Description'] = df['Description'].apply(literal_eval)
之后移除列表中的第一个元素'About this Space':
descriptions['Description'] = descriptions['Description'].str[1:]
现在的问题是,列表中的元素包含\n等空白符,尝试用.strip()、.replace()或lambda函数处理均无效,比如执行以下代码后返回'nan':
descriptions['Description'] = descriptions['Description'].str.replace('\n', '')
请问最简单的解决方法是什么?
解决方法
因为Description列转换后是列表类型,不再是字符串,所以不能直接使用Pandas的str.xxx系列方法(这类方法仅针对字符串列有效),需要对列表内的每个元素单独处理:
方式1:清理列表内元素,保留列表结构
# 遍历每个列表,对元素先去除首尾空白,再替换换行符 descriptions['Description'] = descriptions['Description'].apply( lambda lst: [item.strip().replace('\n', '') for item in lst] )
方式2:清理后合并为单个字符串(若无需保留列表结构)
如果最终需要的是完整的描述文本而非列表,可以直接将清理后的元素拼接成字符串:
# 清理每个元素后用空格连接,生成完整描述 descriptions['Description'] = descriptions['Description'].apply( lambda lst: ' '.join([item.strip().replace('\n', '') for item in lst]) )
补充说明
之前的str.replace无效是因为:转换列表后,Description列的数据类型是object(存储Python列表),而str.replace是Pandas为字符串列设计的矢量化方法,无法识别列表类型,因此返回NaN。
内容的提问来源于stack exchange,提问作者Adam Idris
相关产品推荐
相关产品推荐

