You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas列中为每行保留唯一值(重复项仅存一次)?

处理Pandas DataFrame每行单词去重(保留首次出现顺序)

解决方案步骤:

假设你的DataFrame中cars列的元素是列表类型,可以直接用以下方法处理:

  1. 导入依赖并构造示例DataFrame(如果已有DataFrame可跳过此步)
import pandas as pd

data = {
    'cars': [
        ['honda', 'toyota'],
        ['honda', 'none', 'honda', 'toyota', 'toyota'],
        ['lexus', 'mazda'],
        ['honda', 'mazda', 'lexus', 'mazda', 'honda']
    ]
}
df = pd.DataFrame(data)
  1. 定义去重函数,保留单词首次出现的顺序
def deduplicate_words(lst):
    seen = set()
    unique_words = []
    for word in lst:
        if word not in seen:
            seen.add(word)
            unique_words.append(word)
    return ', '.join(unique_words)
  1. 将函数应用到cars列
df['cars'] = df['cars'].apply(deduplicate_words)

执行后,DataFrame的cars列就会变成你期望的格式:

cars
0        honda, toyota
1  honda, none, toyota
2        lexus, mazda
3  honda, mazda, lexus

特殊情况处理:如果cars列是字符串格式(比如带方括号的字符串)

如果你的cars列存储的是类似"[honda, toyota]"的字符串,需要先解析成列表再去重:

# 先把字符串转成列表
df['cars'] = df['cars'].str.strip('[]').str.split(', ')
# 再应用去重函数
df['cars'] = df['cars'].apply(deduplicate_words)

内容的提问来源于stack exchange,提问作者CPDatascience

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 11:55:16