You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于user_id和meet_id分组Pandas DataFrame,按频次拼接label字符串

Solution for Label Concatenation by (user_id, meet_id) Group

Let's fix this properly—we need to group by both user_id and meet_id, then handle label concatenation based on frequency rank without splitting label strings. Here's a step-by-step solution:

Step 1: Requirement Recap

For each unique (user_id, meet_id) group:

  • Highest frequency label: Keep the original value
  • Second highest frequency label: Concatenate with the highest frequency label
  • Lower frequency labels: Concatenate with all higher frequency labels (in order from highest to next highest)

Step 2: Full Working Code

import pandas as pd

# Your original DataFrame
mytable = pd.DataFrame(
    {'user_id': ['3c', '3c', '3c', '3c','3c', '3c', '3c', '3c', '3c', '3c', '3c', '3c', '3c', '3d', '3d', '3d', '3d', '3e', '3e', '3r', '3w', '3w', '3w', '3w'],
     'meet_id': [1,1,1,1,1,1,1,2,2,2,2,2,2,3,3,3,3,4,5,6,1,2,1,1],
     'text': ['abc', 'abc', 'abc', 'abc', 'abc', 'abc', 'abc', 'xyz', 'xyz', 'xyz', 'xyz', 'xyz', 'xyz', 'npq', 'npq', 'npq', 'npq', 'tt', 'op', 'li', 'abc', 'xyz', 'abc', 'abc'],
     'label': ['A', 'A', 'A', 'A', 'A','B', 'B', 'B', 'B', 'B', 'C', 'C', 'A', 'G', 'H', 'H', 'H', 'A', 'A', 'B', 'E', 'G', 'B', 'B']}
)
mytable = mytable[['user_id', 'meet_id', 'text', 'label']]

# 1. Calculate label frequencies per (user_id, meet_id) group, sort by frequency descending
label_freq = mytable.groupby(['user_id', 'meet_id'])['label'].value_counts().reset_index(name='count')
sorted_freq = label_freq.sort_values(['user_id', 'meet_id', 'count'], ascending=[True, True, False])

# 2. Build the label-to-concatenated-string mapping for each group
def create_concat_mapping(group):
    sorted_labels = group['label'].tolist()
    label_map = {}
    for idx, lbl in enumerate(sorted_labels):
        if idx == 0:
            # Highest frequency: retain original label
            label_map[lbl] = lbl
        elif idx == 1:
            # Second highest: concat with top label
            label_map[lbl] = f"{lbl}{sorted_labels[0]}"
        else:
            # Lower ranks: concat with all higher-frequency labels (in order)
            label_map[lbl] = f"{lbl}{''.join(sorted_labels[:idx])}"
    return pd.Series(label_map, name='new_label')

# Apply mapping function to each (user_id, meet_id) group
concat_mappings = sorted_freq.groupby(['user_id', 'meet_id']).apply(create_concat_mapping).reset_index()

# 3. Merge the new labels back to the original DataFrame
final_result = mytable.merge(concat_mappings, on=['user_id', 'meet_id', 'label'], how='left')

Step 3: Key Fixes Explained

  • Proper Dual Grouping: Using groupby(['user_id', 'meet_id']) ensures each group is unique to a specific user-meeting pair, fixing the earlier issue of mixing users who shared the same meet_id.
  • No String Splitting: We treat each label as a full string and use ''.join() on the sorted list of labels, preventing accidental character-level splitting of label values.
  • AttributeError Resolution: The custom function and structured grouping avoid invalid operations on grouped objects that caused the earlier error.

Example Output Snippet

For the group user_id='3c', meet_id=2:

  • Original labels and counts: B (4), C (2), A (1)
  • New labels:
    • B → 'B' (highest frequency)
    • C → 'CB' (second highest, concatenated with top label)
    • A → 'ACB' (lowest, concatenated with all higher labels)

内容的提问来源于stack exchange,提问作者Shiva Kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:47:12