如何去除Pandas DataFrame分词后列值的括号格式?
问题:将Pandas中分词后的列表列转为逗号分隔字符串
原始DataFrame定义:
import pandas as pd df = pd.DataFrame({'column1': ['Severe weather Not Severe weather kind of severe weather']})
使用NLTK分词的代码:
from nltk.tokenize import word_tokenize df['column1'] = df['column1'].apply(lambda x: word_tokenize(x))
分词后column1列的显示效果(Pandas以带括号的列表形式展示):
column1
0 [Severe, weather, Not, Severe, weather, kind, of, severe, weather]
期望转换为无括号的逗号分隔格式:
column1
0 Severe, weather, Not, Severe, weather, kind, of, severe, weather
尝试过以下两种无效方法:
方法1:
def delete_brackets(x): for i in x: if i == '[' or i == ']': x.remove(i) return x df=delete_brackets(df)
方法2:
def remove_brackets(x): return x.replace('[', '').replace(']', '') df=remove_brackets(df)
错误原因分析
之前的方法完全没抓准问题本质:
column1中的元素是Python列表对象,显示出来的[]只是Pandas对列表的格式化展示,并非列表里真的包含[或]字符。- 第一个方法遍历的是DataFrame的列名,而非列内的列表元素,逻辑完全错位。
- 第二个方法试图对DataFrame调用字符串replace方法,但replace只适用于字符串类型,对列表对象无效。
正确解决方法
方法1:使用apply结合str.join
直接对列中的每个列表进行字符串拼接:
df['column1'] = df['column1'].apply(lambda x: ', '.join(x))
方法2:使用Pandas矢量化str.join方法
Pandas针对列表类型的列提供了内置的str.join操作,更简洁高效:
df['column1'] = df['column1'].str.join(', ')
执行后,column1列的内容就会变成期望的逗号分隔字符串格式,不再显示括号。
内容的提问来源于stack exchange,提问作者xavi
相关产品推荐
相关产品推荐

