如何使用存储在类字典结构中的1至n个条件清洗Pandas DataFrame?
Great question! You're right that nested np.where() gets unwieldy when you have multiple rules—instead, we can programmatically build boolean masks for each rule and apply them to update your DataFrame. This approach scales seamlessly for 1 to n rules, and keeps your code clean and maintainable.
Step 1: Simplify Your Rule Structure (Optional but Recommended)
First, let's tweak your rules format to make it easier to parse. Instead of nested lists, use a list of dictionaries where each entry clearly defines the conditions (all ANDed together) and the columns to update:
rules = [ { "conditions": {"shape": "round", "color": "blue"}, "results": {"fruit": "blueberry"} }, { "conditions": {"fruit": "orange", "shape": "long", "color": "yellow"}, # Fixed typo: "yeelow" → "yellow" "results": {"fruit": "banana"} }, { "conditions": {"fruit": "apple", "color": "green"}, "results": {"shape": "round"} } ]
Step 2: Programatically Apply Each Rule
We'll loop through each rule, build a boolean mask that checks all conditions (using & for logical AND), then use df.loc[] to update the target columns only where the mask is True.
Here's the full code:
import pandas as pd import numpy as np # Sample DataFrame d = [ ["apple", "square", "green"], ["apple", "round", "blue"], ["orange", "long", "yellow"], ] df = pd.DataFrame(d, columns=["fruit", "shape", "color"]) # Simplified rules rules = [ { "conditions": {"shape": "round", "color": "blue"}, "results": {"fruit": "blueberry"} }, { "conditions": {"fruit": "orange", "shape": "long", "color": "yellow"}, "results": {"fruit": "banana"} }, { "conditions": {"fruit": "apple", "color": "green"}, "results": {"shape": "round"} } ] # Apply each rule for rule in rules: # Build the boolean mask: all conditions must be true (AND) mask = pd.Series([True]*len(df), index=df.index) for col, val in rule["conditions"].items(): mask &= df[col] == val # Update each target column with the new value for target_col, new_val in rule["results"].items(): df.loc[mask, target_col] = new_val print(df)
Output:
fruit shape color 0 apple round green 1 blueberry round blue 2 banana long yellow
Step 3: Working With Your Original Rule Format
If you need to stick with your original nested rules structure, we just need to add a small parsing step to extract the conditions and results:
# Your original rule structure original_rules = [ [ { "condition": [{"shape": "round", "color": "blue"}], "result": [{"fruit": "blueberry"}], } ], [ { "condition": [{"fruit": "orange", "shape": "long"}], "result": [{"fruit": "banana"}], } ], [ { "condition": [{"fruit": "apple", "color": "green"}], "result": [{"shape": "round"}], } ], ] # Parse and apply original rules for rule_group in original_rules: for rule in rule_group: # Extract condition and result dictionaries cond_dict = rule["condition"][0] res_dict = rule["result"][0] # Build mask mask = pd.Series([True]*len(df), index=df.index) for col, val in cond_dict.items(): mask &= df[col] == val # Update columns for target_col, new_val in res_dict.items(): df.loc[mask, target_col] = new_val
Why This Works Better Than Nested np.where()
- Scalability: Add/remove rules without rewriting complex nested logic.
- Readability: Each rule is self-contained, making it easy to debug or modify.
- Flexibility: Handle updates to multiple columns per rule (if needed), not just one.
Quick Notes
- Rule order matters: Later rules will overwrite changes from earlier ones if they share overlapping conditions. Adjust the order of your rules list to prioritize updates as needed.
- If you need OR conditions instead of AND, replace
mask &= ...withmask |= ...(logical OR).
内容的提问来源于stack exchange,提问作者PiEnthusiast

