You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何遍历多个DataFrame批量计算分组统计量并生成对应结果变量

How to Batch Process Multiple DataFrames and Generate Corresponding Statistic Series

Hey there! Let's break down how to solve your problem. The "can't assign to operator" error you're seeing comes from trying to use a dynamic expression (like df+'regmedian') as a variable name—Python doesn't let you do that because variable names need to be valid, static identifiers. Here are two ways to achieve what you want, with a recommended best practice first:

Storing your statistical results in a dictionary is the cleanest, most maintainable approach. It keeps all your grouped stats organized in one place, avoids cluttering your namespace with lots of similar variable names, and makes it easy to access or iterate over results later.

Step-by-Step Implementation

First, pair your DataFrames with their names so you can reference them when naming your stats:

# Map DataFrame names to their corresponding objects
df_mapping = {"nw15": nw15, "nw16": nw16, "nw17": nw17}
# Initialize an empty dictionary to hold your results
stat_results = {}

Then loop through the mapping to calculate your desired stats (we'll use median as an example, but you can extend this to mean, geometric mean, etc.):

for df_name, df in df_mapping.items():
    # Calculate grouped median and store it in the dictionary
    stat_results[f"regmedian_{df_name}"] = pd.Series(
        df.groupby(["GORegion"])["PurchasePrice"].median(),
        name=f"regmedian_{df_name}"
    )

To access the median result for nw15, just use stat_results["regmedian_nw15"]—simple and organized!

Extend to Multiple Statistics

If you want to calculate mean, geometric mean, and median all at once, you can generate a DataFrame of stats for each original DataFrame:

from scipy.stats import gmean  # Needed for geometric mean calculation

multi_stat_results = {}
for df_name, df in df_mapping.items():
    grouped_data = df.groupby(["GORegion"])["PurchasePrice"]
    # Create a DataFrame with all desired stats
    multi_stat_results[df_name] = pd.DataFrame({
        "mean": grouped_data.mean(),
        "geometric_mean": grouped_data.apply(gmean),
        "median": grouped_data.median()
    })

Now multi_stat_results["nw15"] will give you a DataFrame with three columns of grouped stats for nw15.

If you absolutely need standalone variables like regmedian_nw15, you can use Python's globals() dictionary to create them dynamically. However, this approach is not ideal because it pollutes your global namespace and makes it harder to track variables as your code grows.

Here's how you'd do it:

# List of tuples with DataFrame names and objects
df_list = [("nw15", nw15), ("nw16", nw16), ("nw17", nw17)]

for df_name, df in df_list:
    # Dynamically create a variable in the global namespace
    globals()[f"regmedian_{df_name}"] = pd.Series(
        df.groupby(["GORegion"])["PurchasePrice"].median(),
        name=f"regmedian_{df_name}"
    )

After running this, you'll have the variables regmedian_nw15, regmedian_nw16, and regmedian_nw17 available in your environment.


内容的提问来源于stack exchange,提问作者Kat Christiansen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:48:36