如何基于Col2的频次值扩展目标DataFrame?
Solution for Expanding DataFrame Rows Based on Column Values
Here's a straightforward approach using pandas to achieve your desired row expansion:
import pandas as pd # Original DataFrame df = pd.DataFrame({ 'Col1': ['a', 'a', 'b', 'b'], 'Col2': [2, 1, 4, 2], 'Col3': [6, 3, 2, 3] }) # Expand rows based on Col2 values expanded_df = df.repeat(df['Col2'])[['Col1', 'Col3']].reset_index(drop=True) print(expanded_df)
Output:
Col1 Col3 0 a 6 1 a 6 2 a 3 3 b 2 4 b 2 5 b 2 6 b 2 7 b 3 8 b 3
How it works:
df.repeat(df['Col2']): This method repeats each row in the DataFrame exactly the number of times specified inCol2for that row. For example, the first row (withCol2=2) is duplicated twice, the second row (Col2=1) stays as a single row, etc.[['Col1', 'Col3']]: We select only the columns needed in the final output, droppingCol2since it's no longer required.reset_index(drop=True): This resets the index of the expanded DataFrame to a clean sequential index (0,1,2,...) instead of retaining the original duplicated indices.
Alternative approach using index.repeat:
If you prefer working with indices directly, you can also use this equivalent method:
expanded_df = df.loc[df.index.repeat(df['Col2'])][['Col1', 'Col3']].reset_index(drop=True)
This works by repeating the row indices according to Col2 values, then using those repeated indices to slice the original DataFrame, followed by selecting the desired columns and resetting the index.
内容的提问来源于stack exchange,提问作者NBC
相关产品推荐
相关产品推荐

