Python中基于另一列统计列表列中值的首次出现次数(含同频次场景处理)
Got it, let's tackle this problem step by step. Your original code works for global first occurrences, but it doesn't account for the nuance of same-no_of_values groups. We need to split the count into two parts: elements that are truly unique to their list (strict) and elements that are new to higher groups but shared with other lists in the same priority group.
First, let's recap your input DataFrame for reference:
import pandas as pd df = pd.DataFrame({ 'value': [['AB','BC','CD','DE','EF','FG','GH','HI'], ['BC','CD','DE','IJ','JK','KL','LM'], ['AB','CD','DE','IJ','JK','GH','HI'], ['AB','CD','DE','MN'], ['C', 'D', 'M'], ['MN','NO'], ['APQ']], 'no_of_values': [8,7,7,4,3,2,1] })
Key Logic to Fix the Issue
We'll process groups of rows with the same no_of_values (in descending order, since your DataFrame is already sorted) and track two things:
- Elements that have already appeared in higher
no_of_valuesgroups. - Elements that are shared across multiple lists within the same
no_of_valuesgroup.
Then, for each list, we'll count:
strict_1st_occurence: Elements that haven't been seen in higher groups AND aren't shared with other lists in the same group.shared_1st_occurences: Elements that haven't been seen in higher groups BUT are shared with other lists in the same group.
Full Code Implementation
# Initialize a set to track elements seen in higher-priority groups seen_in_higher_groups = set() # Add empty columns for our results df['strict_1st_occurence'] = 0 df['shared_1st_occurences'] = 0 # Process each group of same no_of_values (keep descending order) for group_value, group_data in df.groupby('no_of_values', sort=False): # Step 1: Collect all elements in the current group and find shared elements all_group_elements = [] for lst in group_data['value']: all_group_elements.extend(lst) # Count how many times each element appears in the group element_counts = pd.Series(all_group_elements).value_counts() # Elements that appear in >=2 lists of the current group shared_in_group = set(element_counts[element_counts >= 2].index) # Step 2: Calculate strict and shared counts for each row in the group for idx, row in group_data.iterrows(): current_list = row['value'] # Filter out elements already seen in higher groups new_elements = [item for item in current_list if item not in seen_in_higher_groups] # Strict count: new elements not shared with other lists in the group strict_count = sum(1 for item in new_elements if item not in shared_in_group) # Shared count: new elements that are shared with other lists in the group shared_count = sum(1 for item in new_elements if item in shared_in_group) df.loc[idx, 'strict_1st_occurence'] = strict_count df.loc[idx, 'shared_1st_occurences'] = shared_count # Step 3: Add all group elements to the seen set for lower-priority groups seen_in_higher_groups.update(all_group_elements) # View the final result print(df[['value', 'no_of_values', 'strict_1st_occurence', 'shared_1st_occurences']])
Expected Output
This code will produce exactly the result you wanted:
value no_of_values strict_1st_occurence shared_1st_occurences 0 [AB, BC, CD, DE, EF, FG, GH, HI] 8 8 0 1 [BC, CD, DE, IJ, JK, KL, LM] 7 2 2 2 [AB, CD, DE, IJ, JK, GH, HI] 7 0 2 3 [AB, CD, DE, MN] 4 1 0 4 [C, D, M] 3 3 0 5 [MN, NO] 2 1 0 6 [APQ] 1 1 0
Why This Works
- The
seen_in_higher_groupsset ensures we never count elements that appeared in a group with a higherno_of_values. - By analyzing each group's elements, we can distinguish between elements that are unique to a single list (strict) and those that are new but shared across the group.
- Processing groups in descending order (using
sort=Falseto preserve your DataFrame's existing order) ensures we build the seen set correctly.
内容的提问来源于stack exchange,提问作者user18334962
相关产品推荐
相关产品推荐

