在Pandas DataFrame中为随机值匹配最近更大sum值对应行
Got it, let's work through this problem step by step. Here's a straightforward, efficient way to generate your 5000-row DataFrame using pandas and Python's built-in tools:
1. Import Required Libraries
First, we'll pull in the tools we need: pandas for DataFrame operations, numpy to generate random values, and bisect for efficient binary search (since your sum column is monotonically increasing, this will save us a ton of time compared to looping through all rows for each random value).
import pandas as pd import numpy as np import bisect
2. Define Your Original DataFrame
Let's start by setting up the existing DataFrame you provided:
# Original DataFrame with your sample data original_df = pd.DataFrame({ 'Probability': [0.008773, 0.008715, 0.007244, 0.006997], 'sum': [0.008773, 0.017488, 0.024732, 0.031730] })
3. Generate 5000 Random Values Between 0 and 1
Using numpy's rand() function, we can quickly create 5000 random floats in the [0,1) range:
# Generate 5000 random values random_values = np.random.rand(5000)
4. Find the Closest Larger Value in the sum Column
Since your sum column is sorted in ascending order, binary search is the most efficient way to find the first value larger than each random number. Here's how it works:
- Extract the
sumcolumn into a list for faster lookup - Use
bisect_right()to find the insertion point for each random value (this gives us the index of the first element insumthat's larger than the random value) - If a random value is bigger than all elements in
sum, we'll default to the last row (since there's no larger value available)
# Extract sum values into a list for binary search sum_values = original_df['sum'].tolist() # Get indices of the closest larger values selected_indices = [] for val in random_values: idx = bisect.bisect_right(sum_values, val) # Handle cases where no larger value exists if idx == len(sum_values): selected_indices.append(len(sum_values) - 1) else: selected_indices.append(idx)
5. Create the Final 5000-Row DataFrame
Now we'll use the indices we found to pull the corresponding rows from the original DataFrame, then reset the index to clean up the result:
# Generate the new DataFrame new_df = original_df.iloc[selected_indices].reset_index(drop=True) # Optional: Verify the result shape (should be 5000 rows, 2 columns) print(new_df.shape)
Key Notes:
- If your actual
sumcolumn isn't sorted, you'll need to sort it first (and keep track of original indices if you need to preserve the original row order) - The fallback to the last row for values larger than all
sumentries is a practical default. If you want to handle these cases differently (e.g., skip them), you can adjust the loop logic accordingly
内容的提问来源于stack exchange,提问作者Сергей Гульдяшев

