You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Pandas中按分组合并长格式列表元素的实现方法

Solution for Combining Lists by Group in Pandas

Hey there, let's solve this problem where we need to group rows by study_id, combine all non-null list values into a single deduplicated list, and then fill that list into every row of the group. Here's a straightforward way to do this using pandas:

Step 1: Set up your data and imports

First, let's import the necessary libraries and create the sample DataFrame to work with:

import pandas as pd
import numpy as np

# Create the input DataFrame
df = pd.DataFrame({
    'study_id': [1, 1, 1, 2, 2, 2],
    'list_value': [['aaa', 'bbb'], ['aaa'], ['ccc'], ['ddd', 'eee', 'aaa'], np.NaN, ['zzz', 'aaa', 'bbb']]
})

Step 2: Define a function to combine and deduplicate lists

We'll write a helper function that takes a group of list values, filters out any NaN entries, concatenates all the lists, and removes duplicates. We can choose to keep the order of first occurrence or just use a set for unordered deduplication:

def combine_and_deduplicate(group):
    # Flatten all non-null lists in the group
    flattened = []
    for lst in group.dropna():
        flattened.extend(lst)
    
    # Option 1: Keep order of first occurrence (Python 3.7+ dicts are ordered)
    unique_list = list(dict.fromkeys(flattened))
    
    # Option 2: Unordered deduplication (faster, no order guarantee)
    # unique_list = list(set(flattened))
    
    return unique_list

Step 3: Group by study_id and propagate the combined list

Next, we'll compute the combined deduplicated list for each study_id group, then map this result back to every row in the original DataFrame:

# Calculate the combined list for each study_id
grouped_results = df.groupby('study_id')['list_value'].apply(combine_and_deduplicate)

# Map the results to each row in the original DataFrame
df['list_value'] = df['study_id'].map(grouped_results)

Step 4: View the final result

Running the code above will give you exactly the output you're looking for:

study_id                  list_value
0         1        ['aaa', 'bbb', 'ccc']
1         1        ['aaa', 'bbb', 'ccc']
2         1        ['aaa', 'bbb', 'ccc']
3         2  ['ddd', 'eee', 'aaa', 'zzz', 'bbb']
4         2  ['ddd', 'eee', 'aaa', 'zzz', 'bbb']
5         2  ['ddd', 'eee', 'aaa', 'zzz', 'bbb']

Key Notes

  • The dropna() call ensures we ignore any np.NaN entries in the list_value column when combining lists.
  • Choose between ordered or unordered deduplication based on your needs—ordered uses dict.fromkeys() to preserve the first occurrence order, while unordered uses a set for simplicity.

内容的提问来源于stack exchange,提问作者KubiK888

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:08:27