基于user_id和meet_id分组Pandas DataFrame,按频次拼接label字符串
Solution for Label Concatenation by (user_id, meet_id) Group
Let's fix this properly—we need to group by both user_id and meet_id, then handle label concatenation based on frequency rank without splitting label strings. Here's a step-by-step solution:
Step 1: Requirement Recap
For each unique (user_id, meet_id) group:
- Highest frequency label: Keep the original value
- Second highest frequency label: Concatenate with the highest frequency label
- Lower frequency labels: Concatenate with all higher frequency labels (in order from highest to next highest)
Step 2: Full Working Code
import pandas as pd # Your original DataFrame mytable = pd.DataFrame( {'user_id': ['3c', '3c', '3c', '3c','3c', '3c', '3c', '3c', '3c', '3c', '3c', '3c', '3c', '3d', '3d', '3d', '3d', '3e', '3e', '3r', '3w', '3w', '3w', '3w'], 'meet_id': [1,1,1,1,1,1,1,2,2,2,2,2,2,3,3,3,3,4,5,6,1,2,1,1], 'text': ['abc', 'abc', 'abc', 'abc', 'abc', 'abc', 'abc', 'xyz', 'xyz', 'xyz', 'xyz', 'xyz', 'xyz', 'npq', 'npq', 'npq', 'npq', 'tt', 'op', 'li', 'abc', 'xyz', 'abc', 'abc'], 'label': ['A', 'A', 'A', 'A', 'A','B', 'B', 'B', 'B', 'B', 'C', 'C', 'A', 'G', 'H', 'H', 'H', 'A', 'A', 'B', 'E', 'G', 'B', 'B']} ) mytable = mytable[['user_id', 'meet_id', 'text', 'label']] # 1. Calculate label frequencies per (user_id, meet_id) group, sort by frequency descending label_freq = mytable.groupby(['user_id', 'meet_id'])['label'].value_counts().reset_index(name='count') sorted_freq = label_freq.sort_values(['user_id', 'meet_id', 'count'], ascending=[True, True, False]) # 2. Build the label-to-concatenated-string mapping for each group def create_concat_mapping(group): sorted_labels = group['label'].tolist() label_map = {} for idx, lbl in enumerate(sorted_labels): if idx == 0: # Highest frequency: retain original label label_map[lbl] = lbl elif idx == 1: # Second highest: concat with top label label_map[lbl] = f"{lbl}{sorted_labels[0]}" else: # Lower ranks: concat with all higher-frequency labels (in order) label_map[lbl] = f"{lbl}{''.join(sorted_labels[:idx])}" return pd.Series(label_map, name='new_label') # Apply mapping function to each (user_id, meet_id) group concat_mappings = sorted_freq.groupby(['user_id', 'meet_id']).apply(create_concat_mapping).reset_index() # 3. Merge the new labels back to the original DataFrame final_result = mytable.merge(concat_mappings, on=['user_id', 'meet_id', 'label'], how='left')
Step 3: Key Fixes Explained
- Proper Dual Grouping: Using
groupby(['user_id', 'meet_id'])ensures each group is unique to a specific user-meeting pair, fixing the earlier issue of mixing users who shared the samemeet_id. - No String Splitting: We treat each
labelas a full string and use''.join()on the sorted list of labels, preventing accidental character-level splitting of label values. - AttributeError Resolution: The custom function and structured grouping avoid invalid operations on grouped objects that caused the earlier error.
Example Output Snippet
For the group user_id='3c', meet_id=2:
- Original labels and counts:
B(4),C(2),A(1) - New labels:
B→'B'(highest frequency)C→'CB'(second highest, concatenated with top label)A→'ACB'(lowest, concatenated with all higher labels)
内容的提问来源于stack exchange,提问作者Shiva Kumar
相关产品推荐
相关产品推荐

