Python Pandas中按分组合并长格式列表元素的实现方法
Hey there, let's solve this problem where we need to group rows by study_id, combine all non-null list values into a single deduplicated list, and then fill that list into every row of the group. Here's a straightforward way to do this using pandas:
Step 1: Set up your data and imports
First, let's import the necessary libraries and create the sample DataFrame to work with:
import pandas as pd import numpy as np # Create the input DataFrame df = pd.DataFrame({ 'study_id': [1, 1, 1, 2, 2, 2], 'list_value': [['aaa', 'bbb'], ['aaa'], ['ccc'], ['ddd', 'eee', 'aaa'], np.NaN, ['zzz', 'aaa', 'bbb']] })
Step 2: Define a function to combine and deduplicate lists
We'll write a helper function that takes a group of list values, filters out any NaN entries, concatenates all the lists, and removes duplicates. We can choose to keep the order of first occurrence or just use a set for unordered deduplication:
def combine_and_deduplicate(group): # Flatten all non-null lists in the group flattened = [] for lst in group.dropna(): flattened.extend(lst) # Option 1: Keep order of first occurrence (Python 3.7+ dicts are ordered) unique_list = list(dict.fromkeys(flattened)) # Option 2: Unordered deduplication (faster, no order guarantee) # unique_list = list(set(flattened)) return unique_list
Step 3: Group by study_id and propagate the combined list
Next, we'll compute the combined deduplicated list for each study_id group, then map this result back to every row in the original DataFrame:
# Calculate the combined list for each study_id grouped_results = df.groupby('study_id')['list_value'].apply(combine_and_deduplicate) # Map the results to each row in the original DataFrame df['list_value'] = df['study_id'].map(grouped_results)
Step 4: View the final result
Running the code above will give you exactly the output you're looking for:
study_id list_value 0 1 ['aaa', 'bbb', 'ccc'] 1 1 ['aaa', 'bbb', 'ccc'] 2 1 ['aaa', 'bbb', 'ccc'] 3 2 ['ddd', 'eee', 'aaa', 'zzz', 'bbb'] 4 2 ['ddd', 'eee', 'aaa', 'zzz', 'bbb'] 5 2 ['ddd', 'eee', 'aaa', 'zzz', 'bbb']
Key Notes
- The
dropna()call ensures we ignore anynp.NaNentries in thelist_valuecolumn when combining lists. - Choose between ordered or unordered deduplication based on your needs—ordered uses
dict.fromkeys()to preserve the first occurrence order, while unordered uses asetfor simplicity.
内容的提问来源于stack exchange,提问作者KubiK888

