基于列名与值的DataFrame切片及多维度分组列表生成需求
Hey there! Let's figure out how to generate those language-Account value lists you need. Using pandas, this task is pretty straightforward—here's a step-by-step breakdown:
1. Set up your sample DataFrame
First, let's recreate the example data you provided to test our solution:
import pandas as pd # Sample data matching your example data = { "EN": ["Milan", "Florence", "London", "Belgrade"], "DE": ["Mailand", "Florenz", "London", "Belgrad"], "IT": ["Milano", "Firenze", "Londra", "Belgrado"], "Account": ["Italy", "Italy", "UK", "World"] } df = pd.DataFrame(data)
2. Generate the language-Account combination lists
We'll use groupby to cluster rows by the Account column, then collect the values from each language column into lists. We'll store these in a dictionary where keys follow the format {Language}_{Account}:
Option 1: Readable step-by-step code
This version is great for understanding what's happening at each stage:
# Initialize an empty dictionary to hold our results result_dict = {} # Loop through each language column (all columns except the last 'Account' column) for lang_col in df.columns[:-1]: # Group rows by Account and aggregate the language column values into lists grouped_data = df.groupby("Account")[lang_col].apply(list) # Populate the dictionary with keys like "EN_Italy" and their corresponding lists for account_name, city_list in grouped_data.items(): result_key = f"{lang_col}_{account_name}" result_dict[result_key] = city_list
Option 2: Concise dictionary comprehension
If you prefer a more compact one-liner (functionally identical to the above):
result_dict = { f"{lang}_{acc}": vals for lang in df.columns[:-1] for acc, vals in df.groupby("Account")[lang].apply(list).items() }
3. Verify the output
To check that we got the desired result, print the dictionary:
for key, value in result_dict.items(): print(f"{key} = {value}")
This will output exactly what you're looking for:
EN_Italy = ['Milan', 'Florence'] EN_UK = ['London'] EN_World = ['Belgrade'] DE_Italy = ['Mailand', 'Florenz'] DE_UK = ['London'] DE_World = ['Belgrad'] IT_Italy = ['Milano', 'Firenze'] IT_UK = ['Londra'] IT_World = ['Belgrado']
Key notes
- The code works for any number of language columns (not just EN/DE/IT) as long as
Accountis the last column. - If your
Accountcolumn isn't the last one, you can adjust the loop to exclude it explicitly:[col for col in df.columns if col != "Account"]instead ofdf.columns[:-1].
内容的提问来源于stack exchange,提问作者Roberto Bertinetti

