基于PAGE_NAME列拆分Pandas DataFrame并设置对应变量名
Hey there! Let's fix your issue with splitting your original DataFrame into individual named DataFrames for each unique PAGE_NAME value. Your current function generates a list of DataFrames, but we can adjust this to either store them in a manageable dictionary (the recommended approach) or create standalone variables directly.
Approach 1: Store in a Dictionary (Recommended)
Using a dictionary is way better for managing multiple DataFrames—it keeps everything organized, makes it easy to iterate over subsets, and avoids cluttering your namespace with 8 separate variables. Here's how to modify your function:
def make_pagename_dataframes(page_name_list, original_df): page_df_dict = {} for page_name in page_name_list: # Create the variable-style key (e.g., "Demographics_df") df_key = f"{page_name}_df" # Use .copy() to avoid unintended view/link to the original DataFrame page_df_dict[df_key] = original_df.loc[original_df['PAGE_NAME'] == page_name].copy() return page_df_dict # Generate the dictionary of DataFrames page_dfs = make_pagename_dataframes(my_list_of_strings, original_df)
To access a specific subset later, just use the key:
# Get the Demographics subset demographics_data = page_dfs['Demographics_df'] # Or use it directly for operations page_dfs['LeadingCausesOfDeath_df'].describe()
This approach is far more maintainable—if you ever add more PAGE_NAME values, you won't have to manually create new variables.
Approach 2: Create Standalone Variables (Not Recommended)
If you absolutely need standalone variables like Demographics_df, you can dynamically create them using Python's globals() function. Note that this makes your code less readable (variables appear out of nowhere) and harder to debug, but here's how to do it:
def make_pagename_dataframes(page_name_list, original_df): for page_name in page_name_list: var_name = f"{page_name}_df" # Create a global variable with the desired name globals()[var_name] = original_df.loc[original_df['PAGE_NAME'] == page_name].copy() # Run the function to create the variables make_pagename_dataframes(my_list_of_strings, original_df) # Now you can use the variables directly print(SummaryMeasuresOfHealth_df.head())
A Quick Note on .copy()
I added .copy() to both examples because Pandas returns a view of the original DataFrame when slicing, not a full copy. Without .copy(), changes to your subset DataFrames could accidentally modify the original data—this avoids that issue.
Why Avoid the Original List Approach?
Your current function returns a list, but extracting subsets from it requires manually matching indices to PAGE_NAME values, which is error-prone:
# Risky—easy to mix up the order of items in the list Demographics_df = list_of_new_dfs[0] SummaryMeasuresOfHealth_df = list_of_new_dfs[1]
If the order of my_list_of_strings ever changes, your variables will point to the wrong data. Stick with the dictionary or dynamic variable approach instead.
内容的提问来源于stack exchange,提问作者TJE

