You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PAGE_NAME列拆分Pandas DataFrame并设置对应变量名

Split DataFrame into Named Subsets by PAGE_NAME

Hey there! Let's fix your issue with splitting your original DataFrame into individual named DataFrames for each unique PAGE_NAME value. Your current function generates a list of DataFrames, but we can adjust this to either store them in a manageable dictionary (the recommended approach) or create standalone variables directly.

Using a dictionary is way better for managing multiple DataFrames—it keeps everything organized, makes it easy to iterate over subsets, and avoids cluttering your namespace with 8 separate variables. Here's how to modify your function:

def make_pagename_dataframes(page_name_list, original_df):
    page_df_dict = {}
    for page_name in page_name_list:
        # Create the variable-style key (e.g., "Demographics_df")
        df_key = f"{page_name}_df"
        # Use .copy() to avoid unintended view/link to the original DataFrame
        page_df_dict[df_key] = original_df.loc[original_df['PAGE_NAME'] == page_name].copy()
    return page_df_dict

# Generate the dictionary of DataFrames
page_dfs = make_pagename_dataframes(my_list_of_strings, original_df)

To access a specific subset later, just use the key:

# Get the Demographics subset
demographics_data = page_dfs['Demographics_df']
# Or use it directly for operations
page_dfs['LeadingCausesOfDeath_df'].describe()

This approach is far more maintainable—if you ever add more PAGE_NAME values, you won't have to manually create new variables.

If you absolutely need standalone variables like Demographics_df, you can dynamically create them using Python's globals() function. Note that this makes your code less readable (variables appear out of nowhere) and harder to debug, but here's how to do it:

def make_pagename_dataframes(page_name_list, original_df):
    for page_name in page_name_list:
        var_name = f"{page_name}_df"
        # Create a global variable with the desired name
        globals()[var_name] = original_df.loc[original_df['PAGE_NAME'] == page_name].copy()

# Run the function to create the variables
make_pagename_dataframes(my_list_of_strings, original_df)

# Now you can use the variables directly
print(SummaryMeasuresOfHealth_df.head())

A Quick Note on .copy()

I added .copy() to both examples because Pandas returns a view of the original DataFrame when slicing, not a full copy. Without .copy(), changes to your subset DataFrames could accidentally modify the original data—this avoids that issue.

Why Avoid the Original List Approach?

Your current function returns a list, but extracting subsets from it requires manually matching indices to PAGE_NAME values, which is error-prone:

# Risky—easy to mix up the order of items in the list
Demographics_df = list_of_new_dfs[0]
SummaryMeasuresOfHealth_df = list_of_new_dfs[1]

If the order of my_list_of_strings ever changes, your variables will point to the wrong data. Stick with the dictionary or dynamic variable approach instead.

内容的提问来源于stack exchange,提问作者TJE

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:44:01