如何高效为Pandas DataFrame添加‘主餐’列:识别食材重叠食谱
Efficient Solution for Assigning Master Meals in Pandas
To solve this problem efficiently (avoiding the O(n²) overhead of apply()), we can use a dictionary-based approach to track valid candidate recipes and their ingredient pairs, allowing O(1) lookups for each row. Here's how it works:
Approach
- Precompute Recipe Indices: Create a mapping from recipe names to their row indices for quick lookup of the earliest candidate.
- Track Valid Candidates: Use a dictionary to store the earliest valid recipe for each possible pair of ingredients (vegetable-fruit, vegetable-protein, fruit-protein). Valid candidates are recipes that either:
- Are their own master (first row), or
- Have no assigned master (so they can be masters for future rows).
- Iterate Through Rows: For each row:
- Check if any of its ingredient pairs exist in the candidate dictionary.
- If matches are found, select the earliest candidate as the master.
- If no matches, mark the master as
Noneand add the row's pairs to the dictionary (if not already present) to be a candidate for future rows.
Code Implementation
import pandas as pd # Example DataFrame df = pd.DataFrame({ 'recipe': ['meal 1', 'meal 2', 'meal 3', 'meal 4', 'meal 5'], 'vegetable': ['carrot', 'carrot', 'beets', 'carrot', 'artichoke'], 'fruit': ['banana', 'apple', 'banana', 'banana', 'banana'], 'protein': ['beef', 'chicken', 'beef', 'fish', 'fish'], 'calories': [10, 50, 100, 150, 200] }) # Initialize master meal column df['master meal'] = None # Precompute mapping from recipe name to its row index recipe_to_index = df['recipe'].reset_index().set_index('recipe')['index'].to_dict() # Dictionary to store earliest valid recipe for each ingredient pair pair_to_recipe = {} # Process first row (its own master) first_idx = 0 first_recipe = df.loc[first_idx, 'recipe'] df.loc[first_idx, 'master meal'] = first_recipe # Add first row's ingredient pairs to the dictionary veg, fruit, prot = df.loc[first_idx, ['vegetable', 'fruit', 'protein']] pair_to_recipe[(veg, fruit)] = first_recipe pair_to_recipe[(veg, prot)] = first_recipe pair_to_recipe[(fruit, prot)] = first_recipe # Iterate over remaining rows for i in range(1, len(df)): current_recipe = df.loc[i, 'recipe'] veg, fruit, prot = df.loc[i, ['vegetable', 'fruit', 'protein']] # Generate all three ingredient pairs for the current row current_pairs = [(veg, fruit), (veg, prot), (fruit, prot)] # Collect all valid candidate recipes from the pair dictionary candidates = set() for p in current_pairs: if p in pair_to_recipe: candidates.add(pair_to_recipe[p]) if candidates: # Find the earliest candidate by its row index candidate_indices = [recipe_to_index[rec] for rec in candidates] earliest_idx = min(candidate_indices) df.loc[i, 'master meal'] = df.loc[earliest_idx, 'recipe'] else: df.loc[i, 'master meal'] = None # If current row has no master, add its pairs to the dictionary (only if not already present) if df.loc[i, 'master meal'] is None: for p in current_pairs: if p not in pair_to_recipe: pair_to_recipe[p] = current_recipe print(df)
Output
recipe vegetable fruit protein calories master meal 0 meal 1 carrot banana beef 10 meal 1 1 meal 2 carrot apple chicken 50 None 2 meal 3 beets banana beef 100 meal 1 3 meal 4 carrot banana fish 150 meal 1 4 meal 5 artichoke banana fish 200 None
Efficiency
This approach runs in O(n) time complexity, where n is the number of rows. Each row is processed exactly once, with constant-time operations (dictionary lookups and inserts) for each row. This is drastically faster than the O(n²) apply() method for large datasets, as it avoids comparing every row to all previous rows.
内容的提问来源于stack exchange,提问作者Malisz
相关产品推荐
相关产品推荐

