如何在Pandas列中为每行保留唯一值(重复项仅存一次)?
处理Pandas DataFrame每行单词去重(保留首次出现顺序)
解决方案步骤:
假设你的DataFrame中cars列的元素是列表类型,可以直接用以下方法处理:
- 导入依赖并构造示例DataFrame(如果已有DataFrame可跳过此步)
import pandas as pd data = { 'cars': [ ['honda', 'toyota'], ['honda', 'none', 'honda', 'toyota', 'toyota'], ['lexus', 'mazda'], ['honda', 'mazda', 'lexus', 'mazda', 'honda'] ] } df = pd.DataFrame(data)
- 定义去重函数,保留单词首次出现的顺序
def deduplicate_words(lst): seen = set() unique_words = [] for word in lst: if word not in seen: seen.add(word) unique_words.append(word) return ', '.join(unique_words)
- 将函数应用到
cars列
df['cars'] = df['cars'].apply(deduplicate_words)
执行后,DataFrame的cars列就会变成你期望的格式:
cars 0 honda, toyota 1 honda, none, toyota 2 lexus, mazda 3 honda, mazda, lexus
特殊情况处理:如果cars列是字符串格式(比如带方括号的字符串)
如果你的cars列存储的是类似"[honda, toyota]"的字符串,需要先解析成列表再去重:
# 先把字符串转成列表 df['cars'] = df['cars'].str.strip('[]').str.split(', ') # 再应用去重函数 df['cars'] = df['cars'].apply(deduplicate_words)
内容的提问来源于stack exchange,提问作者CPDatascience
相关产品推荐
相关产品推荐

