使用Pandas根据其他列添加随机值:同ID组NPS值一致需求
How to Add a Consistent NPS Column to Your Ad DataFrame
Hey Alina, I totally get where you're coming from—using a dictionary to map ID combinations to NPS makes sense, and I’ll walk you through a couple of straightforward ways to pull this off with pandas. Let’s dive in!
Method 1: Use groupby + transform (Most Concise)
This is the cleanest approach. We’ll group your DataFrame by the three ID columns, generate a single random NPS value for each group, then broadcast that value to every row in the group.
import pandas as pd import numpy as np # Example DataFrame (replace with your actual data) df = pd.DataFrame({ 'OfferID': [1, 1, 2, 2, 3, 3, 3], 'SiteID': [10, 10, 20, 20, 30, 30, 30], 'CategoryID': [100, 100, 200, 200, 300, 300, 300] }) # Add NPS column with consistent values per ID combination df['NPS'] = df.groupby(['OfferID', 'SiteID', 'CategoryID'])['OfferID'].transform( lambda group: np.random.randint(1, 11) # Generates random int from 1-10 inclusive )
How it works:
groupby(['OfferID', 'SiteID', 'CategoryID'])clusters rows that share all three IDs together.transformapplies the lambda function to each group, then copies the resulting value to every row in that group. This ensures identical ID combinations get the same NPS.
Method 2: Create a Mapping Dictionary (Your Original Idea!)
If you prefer to explicitly build the ID-to-NPS mapping (great for transparency), here’s how to implement it:
# Step 1: Get all unique ID combinations unique_groups = df[['OfferID', 'SiteID', 'CategoryID']].drop_duplicates().reset_index(drop=True) # Step 2: Assign random NPS to each unique combination unique_groups['NPS'] = np.random.randint(1, 11, size=len(unique_groups)) # Step 3: Merge the mapping back into your original DataFrame df = df.merge(unique_groups, on=['OfferID', 'SiteID', 'CategoryID'], how='left')
How it works:
- We first extract only the unique ID pairs (no duplicate rows) to avoid generating redundant NPS values.
- We generate a random 1-10 value for each unique group, then merge this mapping back into the original DataFrame. Every row with matching IDs will inherit the same NPS.
Bonus: Make Results Reproducible
If you want the same random NPS values every time you run the code, set a random seed before generating numbers:
np.random.seed(42) # Use any integer you like # Then run either method above
内容的提问来源于stack exchange,提问作者Alina
相关产品推荐
相关产品推荐

