如何通过循环获取Pandas DataFrame分组对应的唯一Colour值
解决方案
方法一:使用Pandas原生groupby(推荐)
直接按Name和Age分组,对Colour列取唯一值并转为列表,这是最贴合Pandas风格且高效的实现方式:
import pandas as pd data = {'Name':['Mathew', 'Mathew', 'Mathew', 'Mathew','Mathew','John','John','John'], 'Age':[12,12, 12,13, 13,12,13,13], 'Colour':['Yellow','Blue','Yellow','green','blue','pink','black','brown']} df = pd.DataFrame(data) # 分组并提取唯一Colour列表 grouped_result = df.groupby(['Name', 'Age'])['Colour'].unique().apply(list) # 遍历输出结果 for (name, age), colour_list in grouped_result.items(): print(f"Name: {name}, Age: {age} -> 唯一Colour列表: {colour_list}")
输出结果:
Name: John, Age: 12 -> 唯一Colour列表: ['pink'] Name: John, Age: 13 -> 唯一Colour列表: ['black', 'brown'] Name: Mathew, Age: 12 -> 唯一Colour列表: ['Yellow', 'Blue'] Name: Mathew, Age: 13 -> 唯一Colour列表: ['green', 'blue']
方法二:改进你的嵌套循环
如果坚持用循环实现,需要在每次循环中过滤出当前Name和Age对应的行,再提取唯一的Colour值:
import pandas as pd data = {'Name':['Mathew', 'Mathew', 'Mathew', 'Mathew','Mathew','John','John','John'], 'Age':[12,12, 12,13, 13,12,13,13], 'Colour':['Yellow','Blue','Yellow','green','blue','pink','black','brown']} df = pd.DataFrame(data) # 遍历所有唯一的Name和Age组合 for name in df['Name'].unique(): for age in df['Age'].unique(): # 过滤当前Name和Age对应的行 filtered_rows = df[(df['Name'] == name) & (df['Age'] == age)] # 仅当该组合存在数据时处理 if not filtered_rows.empty: unique_colours = filtered_rows['Colour'].unique().tolist() print(f"Age: {age}, Name: {name} -> 唯一Colour列表: {unique_colours}")
输出结果:
Age: 12, Name: Mathew -> 唯一Colour列表: ['Yellow', 'Blue'] Age: 13, Name: Mathew -> 唯一Colour列表: ['green', 'blue'] Age: 12, Name: John -> 唯一Colour列表: ['pink'] Age: 13, Name: John -> 唯一Colour列表: ['black', 'brown']
注:原代码用set(df.Name)会打乱顺序,改用df['Name'].unique()可以保留数据中原本的出现顺序。
内容的提问来源于stack exchange,提问作者Mathew John
相关产品推荐
相关产品推荐

