如何将DataFrame修改为每行仅保留一个非零值其余均为0(随机选择)
Hey there! Let's tackle this problem—turning your DataFrame into one where each row has exactly one non-zero value (in a random column) and zeros everywhere else. Here's how to do it smoothly with pandas and numpy:
We'll cover two common scenarios depending on whether you want to keep original values or use custom ones, plus a fast vectorized version for large datasets.
Scenario 1: Keep Original Data's Random Column Value
If you want to retain the existing value from a randomly selected column in each row (and set all others to 0):
import pandas as pd import numpy as np # Replace this with your actual DataFrame df = pd.DataFrame(np.random.randint(1, 100, size=(5, 4)), columns=["A", "B", "C", "D"]) # Pick a random column for each row random_col_indices = np.random.choice(df.columns, size=len(df)) # Create a blank all-zero DataFrame, then fill in the random values result_df = pd.DataFrame(0, index=df.index, columns=df.columns) for row_idx, col in enumerate(random_col_indices): result_df.loc[row_idx, col] = df.loc[row_idx, col] print(result_df)
Scenario 2: Use Custom Non-Zero Values (Like Your Example)
If you want to assign custom values (e.g., 1, 3, 25 as in your sample) to the random columns:
import pandas as pd import numpy as np # Replace with your real DataFrame df = pd.DataFrame(np.random.randint(1, 100, size=(5, 4)), columns=["A", "B", "C", "D"]) # List of custom values (must match the number of rows in your DataFrame) custom_non_zero_values = [1, 3, 25, 7, 42] # Pick random columns for each row random_col_indices = np.random.choice(df.columns, size=len(df)) # Build the result DataFrame result_df = pd.DataFrame(0, index=df.index, columns=df.columns) for row_idx, (col, val) in enumerate(zip(random_col_indices, custom_non_zero_values)): result_df.loc[row_idx, col] = val print(result_df)
Fast Vectorized Version (For Large Datasets)
If you're working with a huge DataFrame (thousands/millions of rows), loops can be slow. Use this vectorized approach instead—it's way faster:
import pandas as pd import numpy as np df = pd.DataFrame(np.random.randint(1, 100, size=(10000, 4)), columns=["A", "B", "C", "D"]) # Generate a mask where only the random column per row is 1, others 0 random_cols = np.random.choice(df.columns, size=len(df)) mask = pd.get_dummies(random_cols).set_index(df.index) # Multiply original DataFrame by mask to keep only random column values result_df = df * mask
Quick Breakdown of Key Bits:
np.random.choice(df.columns, size=len(df)): Picks a random column for every row—this is how we get the random non-zero position per row.- The all-zero starting DataFrame ensures we don't accidentally modify your original data (always a good practice!).
- The vectorized mask method avoids loops entirely, making it perfect for big datasets where speed matters.
内容的提问来源于stack exchange,提问作者mikbo

