You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分组后统计Class_par_ratio频率及最大频率的实现问题

Pandas分组统计频率并添加最大频率列解决方案

Let's fix this step by step, since you're working with Pandas 0.20 and Python 3.6.3, we'll make sure the code is fully compatible with your version.

问题回顾

Your goal is to take a DataFrame with columns FileName, PageNo, LineNo, Name, Class_par_ratio, and:

  • Group by FileName and Class_par_ratio to calculate frequency (stored in Frequency column)
  • Add a Max Freq. column that shows the highest frequency value per FileName group
  • Add a Max_Class column that shows which Class_par_ratio corresponds to that maximum frequency

Your previous attempts fell short:

  1. The first code snippet df.groupby(['FileName'])['Class_par_ratio'].value_counts() only generated basic frequency stats, but didn't add the required max frequency columns or format the output correctly.
  2. The second code had convoluted grouping logic—repeating the same groupby and using agg({'count': max}) was redundant (each group was already unique), and nlargest(1) only kept the top row per FileName, losing all other class data.

Correct Solution (Pandas 0.20 Compatible)

We'll break this into 3 clear steps:

1. Calculate Base Frequencies

First, count occurrences of each (FileName, Class_par_ratio) pair and name the frequency column:

# Count frequencies and reset index to get a flat DataFrame
freq_df = df.groupby(['FileName', 'Class_par_ratio']).size().reset_index(name='Frequency')

For your sample data, this produces:

FileNameClass_par_ratioFrequency
17973375MILK3
17973375OTHER FOODS1
17973375ANIMAL AND VEGETABLE OIL1

2. Add the Max Freq. Column

Use transform('max') to propagate the highest frequency value from each FileName group to every row in that group:

# Attach max frequency per FileName to all rows in the group
freq_df['Max Freq.'] = freq_df.groupby('FileName')['Frequency'].transform('max')

3. Add the Max_Class Column

Find which Class_par_ratio has the maximum frequency per FileName, then merge this back to our main DataFrame:

# Get the Class_par_ratio with the highest frequency per FileName
max_class_df = freq_df.loc[freq_df.groupby('FileName')['Frequency'].idxmax(), ['FileName', 'Class_par_ratio']]
max_class_df.rename(columns={'Class_par_ratio': 'Max_Class'}, inplace=True)

# Merge to add Max_Class to all rows in the original frequency DataFrame
result_df = freq_df.merge(max_class_df, on='FileName', how='left')

Full Combined Code

# Step 1: Calculate frequencies
freq_df = df.groupby(['FileName', 'Class_par_ratio']).size().reset_index(name='Frequency')

# Step 2: Add max frequency column
freq_df['Max Freq.'] = freq_df.groupby('FileName')['Frequency'].transform('max')

# Step 3: Get and merge max frequency class
max_class_df = freq_df.loc[freq_df.groupby('FileName')['Frequency'].idxmax(), ['FileName', 'Class_par_ratio']]
max_class_df.rename(columns={'Class_par_ratio': 'Max_Class'}, inplace=True)
result_df = freq_df.merge(max_class_df, on='FileName', how='left')

# Optional: Reorder columns to match your expected output
result_df = result_df[['FileName', 'Class_par_ratio', 'Frequency', 'Max_Class', 'Max Freq.']]

Sample Output (Matching Your Data)

Running this code on your sample data will produce:

FileNameClass_par_ratioFrequencyMax_ClassMax Freq.
17973375MILK3MILK3
17973375OTHER FOODS1MILK3
17973375ANIMAL AND VEGETABLE OIL1MILK3

This fully meets your requirements: it retains all class entries, shows their individual frequencies, and includes the maximum frequency value and corresponding class for each FileName group.

内容的提问来源于stack exchange,提问作者Madhur Yadav

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:48:53