You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

将DataFrame重复值列转为唯一值,其余列转为年度键值对字典

Solution: Deduplicate Unique Column and Convert A/C to Year-Keyed Dictionaries

Got it, let's fix this up efficiently without messy loops! Your goal is to get a unique Unique column (pun intended) while turning columns A and C into dictionaries where each key is the Year and the value is the corresponding entry from A or C. Here's how to do it cleanly with pandas built-in tools:

First, let's recap your input DataFrame for clarity:

import pandas as pd

df = pd.DataFrame({ 
'A': {0: 'a1', 1: 'a2', 2: 'a3', 3: 'a4'}, 
'Unique': {0: 'b1', 1: 'b1', 2: 'b2', 3: 'b2'}, 
'Year': {0: 2017, 1: 2008, 2: 2017, 3: 2008} , 
'C': {0: 'c1', 1: 'c2', 2: 'c3', 3: 'c4'} 
})

Method 1: Using groupby + agg (Concise)

We can group the DataFrame by the Unique column, then aggregate columns A and C into dictionaries by zipping Year with the column values:

result_df = df.groupby('Unique').agg(
    A=lambda x: dict(zip(df.loc[x.index, 'Year'], x)),
    C=lambda x: dict(zip(df.loc[x.index, 'Year'], x))
).reset_index()

Method 2: Using groupby + apply (More Readable)

If you prefer more explicit code, define a helper function to create the dictionaries for each group:

def build_year_dict(group):
    return pd.Series({
        'A': dict(zip(group['Year'], group['A'])),
        'C': dict(zip(group['Year'], group['C']))
    })

result_df = df.groupby('Unique').apply(build_year_dict).reset_index()

What You'll Get

Running either method will give you this output:

Unique                     A                     C
0     b1  {2017: 'a1', 2008: 'a2'}  {2017: 'c1', 2008: 'c2'}
1     b2  {2017: 'a3', 2008: 'a4'}  {2017: 'c3', 2008: 'c4'}

This works because:

  • groupby('Unique') groups all rows with the same Unique value together
  • The aggregation functions (either the lambda or helper function) take each group, pair the Year values with the corresponding A/C entries, and convert them into a dictionary
  • reset_index() brings the Unique column back into the DataFrame as a regular column instead of the index

This approach is way more efficient than manual loops, especially with larger datasets, and keeps your code clean and maintainable.

内容的提问来源于stack exchange,提问作者sachini_rb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:05:33