如何遍历多个DataFrame批量计算分组统计量并生成对应结果变量
Hey there! Let's break down how to solve your problem. The "can't assign to operator" error you're seeing comes from trying to use a dynamic expression (like df+'regmedian') as a variable name—Python doesn't let you do that because variable names need to be valid, static identifiers. Here are two ways to achieve what you want, with a recommended best practice first:
1. Use a Dictionary to Store Results (Recommended)
Storing your statistical results in a dictionary is the cleanest, most maintainable approach. It keeps all your grouped stats organized in one place, avoids cluttering your namespace with lots of similar variable names, and makes it easy to access or iterate over results later.
Step-by-Step Implementation
First, pair your DataFrames with their names so you can reference them when naming your stats:
# Map DataFrame names to their corresponding objects df_mapping = {"nw15": nw15, "nw16": nw16, "nw17": nw17} # Initialize an empty dictionary to hold your results stat_results = {}
Then loop through the mapping to calculate your desired stats (we'll use median as an example, but you can extend this to mean, geometric mean, etc.):
for df_name, df in df_mapping.items(): # Calculate grouped median and store it in the dictionary stat_results[f"regmedian_{df_name}"] = pd.Series( df.groupby(["GORegion"])["PurchasePrice"].median(), name=f"regmedian_{df_name}" )
To access the median result for nw15, just use stat_results["regmedian_nw15"]—simple and organized!
Extend to Multiple Statistics
If you want to calculate mean, geometric mean, and median all at once, you can generate a DataFrame of stats for each original DataFrame:
from scipy.stats import gmean # Needed for geometric mean calculation multi_stat_results = {} for df_name, df in df_mapping.items(): grouped_data = df.groupby(["GORegion"])["PurchasePrice"] # Create a DataFrame with all desired stats multi_stat_results[df_name] = pd.DataFrame({ "mean": grouped_data.mean(), "geometric_mean": grouped_data.apply(gmean), "median": grouped_data.median() })
Now multi_stat_results["nw15"] will give you a DataFrame with three columns of grouped stats for nw15.
2. Dynamically Generate Variable Names (Not Recommended)
If you absolutely need standalone variables like regmedian_nw15, you can use Python's globals() dictionary to create them dynamically. However, this approach is not ideal because it pollutes your global namespace and makes it harder to track variables as your code grows.
Here's how you'd do it:
# List of tuples with DataFrame names and objects df_list = [("nw15", nw15), ("nw16", nw16), ("nw17", nw17)] for df_name, df in df_list: # Dynamically create a variable in the global namespace globals()[f"regmedian_{df_name}"] = pd.Series( df.groupby(["GORegion"])["PurchasePrice"].median(), name=f"regmedian_{df_name}" )
After running this, you'll have the variables regmedian_nw15, regmedian_nw16, and regmedian_nw17 available in your environment.
内容的提问来源于stack exchange,提问作者Kat Christiansen

