基于Pandas实现多DataFrame按MainName与Datetime索引分组取均值
It looks like the error you’re hitting with pd.concat(keys=longnames) stems from mismatched shapes when trying to split the combined index into Main/Sub names later—especially if your DataFrames have varying date ranges. Instead of relying on the keys parameter, let’s take a more straightforward approach that explicitly attaches the MainName to each DataFrame before combining, then groups to compute the mean.
Here’s a step-by-step implementation:
Step 1: Attach MainName to each DataFrame
Loop through your list of DataFrames and corresponding longnames, extract the MainName from each string, and add it as a new column. This ensures every row is explicitly linked to its MainName, avoiding index shape issues.
import pandas as pd # Assume df_list is your list of time series DataFrames # longnames is your list of 'MainName,SubName' strings for df, name_str in zip(df_list, longnames): # Split the longname to isolate MainName main_name = name_str.split(',')[0] df['MainName'] = main_name
Step 2: Combine all DataFrames
Concatenate the modified DataFrames into a single structure. Since each row now has a MainName column, we don’t need to rely on complex index keys here.
combined_df = pd.concat(df_list)
Step 3: Group by MainName and Date to compute mean
Reset the index to turn the Date index into a column (so we can group by it), then calculate the mean for each combination of MainName and Date. Finally, we’ll set the result back to a hierarchical index as you expected.
# Reset index to make Date a column combined_df = combined_df.reset_index().rename(columns={'index': 'Date'}) # Group by MainName and Date, compute mean of Amount(mm) result = combined_df.groupby(['MainName', 'Date'])['Amount(mm)'].mean() # Optional: Convert to a DataFrame if preferred result_df = result.to_frame()
What the Result Looks Like
The final result (or result_df) will have a MultiIndex where the first level is MainName and the second is Date, with the mean Amount(mm) as the value—exactly matching your expected output structure. For example:
Amount(mm) MainName Date MainName1 1900-01-01 0.0 1900-01-02 0.0 1900-01-03 5.1 ... ... MainName2 1900-01-04 8.0 1900-01-05 9.0 ... ...
This approach avoids the shape mismatch error because we’re explicitly linking each row to its MainName before combining, rather than trying to layer it into the index after concatenation. It’s also more readable and easier to debug if you need to adjust grouping logic later.
内容的提问来源于stack exchange,提问作者WDS

