在Pandas中用首次出现的对应Note与Type替换Code列值
Got it, let's tackle this problem using pandas' built-in vectorized operations—no for loops needed! Here's a straightforward approach that leverages groupby and transform to efficiently map each code to its first-seen note and type.
Step 1: Core Concept
We need to create a "lookup" for each code that stores its first occurrence of type and note, then apply this lookup to every row in the original DataFrame. Pandas' transform method is perfect here because it broadcasts group-level values back to every row in the group, keeping the original DataFrame structure intact.
Step 2: Code Implementation
Let's assume your DataFrame is named df. Here's how to execute the solution:
import pandas as pd # Example input DataFrame (matches your use case) data = { "pid": [1, 1, 2, 2, 3], "code": ["A", "A", "B", "B", "C"], "type": ["drug", "drug", "diag", "diag", "drug"], "note": ["alvedon", "ipren", "headache", "migraine", "paracetamol"] } df = pd.DataFrame(data) # Generate unified type and code name (first occurrence's note) for each code df["unified_type"] = df.groupby("code")["type"].transform("first") df["code_name"] = df.groupby("code")["note"].transform("first") # Optional: Replace original code/type columns with unified values df = df.drop(columns=["code", "type"]).rename(columns={ "code_name": "code", "unified_type": "type" }) print(df)
Step 3: Example Output
Running the code above will produce this result, where every code is replaced by its first-seen note, and type is standardized to the first occurrence:
| pid | code | type | note |
|---|---|---|---|
| 1 | alvedon | drug | alvedon |
| 1 | alvedon | drug | ipren |
| 2 | headache | diag | headache |
| 2 | headache | diag | migraine |
| 3 | paracetamol | drug | paracetamol |
How It Works
groupby("code")["type"].transform("first"): Groups the DataFrame bycode, takes the firsttypevalue from each group, and assigns that value to every row in the group. This creates a Series where every row has the first-seentypefor its code.- The same logic applies to
noteto generatecode_name, replacing the original code with its first associated note. - The final step cleans up the columns by replacing the original
codeandtypewith their unified versions (you can skip this if you want to keep the original columns alongside the unified ones).
Why This Beats For Loops
This approach uses pandas' optimized vectorized operations, which are far faster and more memory-efficient than manual for loops—especially when working with large datasets. It also avoids the risk of off-by-one errors or slow iteration common with loop-based solutions.
内容的提问来源于stack exchange,提问作者AnonX

