You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas按FirstName分组并合并唯一Language列表?

解决Pandas按FirstName分组并获取去重Language数组的问题

Hey there! Let's get your Pandas grouping working correctly. The issue with your current code is that when you use list(set(x)) in the apply function, x is a Series where each element is a single-item list (like ['en']). So set(x) is creating a set of those lists instead of extracting the language strings inside them—definitely not what you want!

正确解法1:先展开列表再分组(推荐,更直观)

This approach mirrors the logic of your SQL query: first "unpack" the Language lists into individual rows, then group by FirstName, collect distinct languages, and sort them (just like ORDER BY in SQL):

import pandas as pd

# 你的原始数据
items = [
    {'FirstName': 'David', 'Language': ['en',]},
    {'FirstName': 'David', 'Language': ['fr',]},
    {'FirstName': 'David', 'Language': ['en',]},
    {'FirstName': 'Bob', 'Language': ['en',]}
]

df = pd.DataFrame(items)

# 1. 把Language列的列表拆分成单独的行
exploded_df = df.explode('Language')

# 2. 分组后获取去重的语言,排序后转为列表
result_df = exploded_df.groupby('FirstName')['Language'].agg(lambda x: sorted(x.unique())).reset_index()

# 3. 转成你需要的字典数组格式
final_result = result_df.to_dict('records')
print(final_result)

输出结果:

[{'FirstName': 'Bob', 'Language': ['en']}, {'FirstName': 'David', 'Language': ['en', 'fr']}]

正确解法2:直接在分组时拼接列表并去重

If you prefer to avoid the explode step, you can use a list comprehension to flatten all the Language lists in each group, then deduplicate and sort:

result_df = df.groupby('FirstName')['Language'].agg(
    lambda x: sorted(list(set(lang for sublist in x for lang in sublist)))
).reset_index()

final_result = result_df.to_dict('records')

为什么你的原始代码不对?

Let's break down the problem with your original code:

df.groupby('FirstName')['Language'].apply(lambda x: list(set(x)))

When you group by FirstName, x for the "David" group is a Series like this:

0    [en]
1    [fr]
2    [en]
Name: Language, dtype: object

Using set(x) here creates a set of list objects ({['en'], ['fr']}), not a set of language strings. Converting that to a list gives you something like [['fr'], ['en']]—which is not the flat, deduplicated list you need.

内容的提问来源于stack exchange,提问作者David542

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:43:58