Pandas中匹配含唯一名称ID与重复名称的DataFrame并生成对应重复ID
No problem! To generate corresponding repeated IDs for your DataFrame with duplicate names, you have two simple and effective approaches in Pandas: using merge() or map(). Let’s break them down with clear examples.
First, let’s set up sample DataFrames to mimic your scenario:
import pandas as pd # DataFrame with unique Name-ID pairs df_unique = pd.DataFrame({ 'Name': ['Alice', 'Bob', 'Charlie'], 'ID': [101, 102, 103] }) # DataFrame with repeated names df_repeated = pd.DataFrame({ 'Name': ['Alice', 'Bob', 'Alice', 'Charlie', 'Bob', 'Bob'] })
Method 1: Using merge() (Left Join)
This is the most intuitive method for combining DataFrames. A left join keeps all rows from your repeated names DataFrame and matches the ID from the unique DataFrame wherever names align.
result = df_repeated.merge(df_unique, on='Name', how='left')
Result:
| Name | ID | |
|---|---|---|
| 0 | Alice | 101 |
| 1 | Bob | 102 |
| 2 | Alice | 101 |
| 3 | Charlie | 103 |
| 4 | Bob | 102 |
| 5 | Bob | 102 |
Method 2: Using map()
If you prefer a more concise approach (or have a large dataset where speed matters), create a name-to-ID dictionary from your unique DataFrame and map it to the repeated names.
# Create a dictionary: {Name: ID} name_id_map = df_unique.set_index('Name')['ID'].to_dict() # Map the IDs to df_repeated df_repeated['ID'] = df_repeated['Name'].map(name_id_map)
This gives you the exact same result as the merge method.
Key Notes:
- If there are names in
df_repeatedthat don’t exist indf_unique, the corresponding ID will beNaN. You can handle this withfillna()if needed (e.g.,df_repeated['ID'] = df_repeated['Name'].map(name_id_map).fillna(0)to replace missing IDs with 0). - Ensure your 'Name' columns are clean! Fix whitespace or case mismatches (like 'alice' vs 'Alice') first:
df_unique['Name'] = df_unique['Name'].str.strip().str.lower() df_repeated['Name'] = df_repeated['Name'].str.strip().str.lower()
内容的提问来源于stack exchange,提问作者mehnaz maharin

