如何重复DataFrame行并为id追加索引以保持其唯一性?
id Unique Got it, let's fix this! You're already using np.repeat to duplicate rows, but the id column is ending up with duplicates. Here's a straightforward, efficient way to make sure each id stays unique by appending an index suffix:
Step 1: Start with your original DataFrame
First, let's set up an example to work with (adjust this to match your actual data):
import pandas as pd import numpy as np # Sample DataFrame df = pd.DataFrame({ 'id': ['user_1', 'user_2', 'user_3'], 'score': [85, 92, 78] })
Step 2: Duplicate rows with np.repeat
Use np.repeat to create your duplicated rows (we'll repeat each row 3 times here—change the number to match your needs):
# Repeat each row 3 times repeat_count = 3 repeated_df = df.loc[np.repeat(df.index, repeat_count)].reset_index(drop=True)
At this point, the id column will have duplicates (e.g., user_1 appears 3 times).
Step 3: Make id unique with group-based indexing
Use groupby and cumcount() to add a unique suffix to each duplicated id. This method is efficient even for large DataFrames:
# Append a sequential number to each duplicated id repeated_df['id'] = ( repeated_df['id'] + '_' + (repeated_df.groupby('id').cumcount() + 1).astype(str) )
What the output looks like
Your final DataFrame will have unique ids while keeping all duplicated row data intact:
id score 0 user_1_1 85 1 user_1_2 85 2 user_1_3 85 3 user_2_1 92 4 user_2_2 92 5 user_2_3 92 6 user_3_1 78 7 user_3_2 78 8 user_3_3 78
Customization tips
- If you don't want an underscore, replace
'_'with another separator (like'-') or remove it entirely to getuser_11,user_12, etc. - If your original
idis numeric, you can modify the logic to add a decimal or offset instead (e.g.,repeated_df['id'] = repeated_df['id'] + (repeated_df.groupby('id').cumcount() + 1)/10to get1.1,1.2, etc.).
内容的提问来源于stack exchange,提问作者Prince Bhatti

