Python入门者求助:如何为ID列添加频次列?
Hey there! No worries at all—we all start somewhere 😊
It sounds like you want to calculate both the mean value and the frequency (count of rows) for each combination of ID and Date in your DataFrame. Instead of splitting this into separate steps, you can do it all in one clean groupby operation using pandas' agg() method, which lets you compute multiple statistics at once.
Here's how to do it:
Method 1: Single Groupby with Multiple Aggregations
This is the most efficient approach since it only requires one pass over your data:
import pandas as pd # Replace "your_numeric_column" with the actual name of the column you're calculating the mean for result = dfxyz.groupby(["ID", "Date"]).agg( average_value=("your_numeric_column", "mean"), frequency=("ID", "size") # `size()` counts the number of rows in each group (your frequency) ).reset_index()
agg()lets you define custom output column names (likeaverage_valueandfrequency) paired with the column to compute stats on and the function to use.reset_index()convertsIDandDatefrom groupby indexes back into regular columns, making the result easier to work with.
Method 2: Combine Existing Mean Data with Frequency
If you already have your mean calculation and want to merge the frequency data later, you can do this:
# Your existing mean calculation (with reset_index to keep ID/Date as columns) dfcount1 = dfxyz.groupby(["ID", "Date"]).mean().reset_index() # Calculate frequency for each ID/Date group frequency_df = dfxyz.groupby(["ID", "Date"]).size().reset_index(name="frequency") # Merge the two DataFrames on ID and Date final_result = pd.merge(dfcount1, frequency_df, on=["ID", "Date"])
Just make sure to replace your_numeric_column with the actual name of the column you're taking the mean of (e.g., if it's a "Sales" column, use "Sales" instead).
内容的提问来源于stack exchange,提问作者Mark White

