You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas DataFrame中为随机值匹配最近更大sum值对应行

Got it, let's work through this problem step by step. Here's a straightforward, efficient way to generate your 5000-row DataFrame using pandas and Python's built-in tools:

Solution Steps

1. Import Required Libraries

First, we'll pull in the tools we need: pandas for DataFrame operations, numpy to generate random values, and bisect for efficient binary search (since your sum column is monotonically increasing, this will save us a ton of time compared to looping through all rows for each random value).

import pandas as pd
import numpy as np
import bisect

2. Define Your Original DataFrame

Let's start by setting up the existing DataFrame you provided:

# Original DataFrame with your sample data
original_df = pd.DataFrame({
    'Probability': [0.008773, 0.008715, 0.007244, 0.006997],
    'sum': [0.008773, 0.017488, 0.024732, 0.031730]
})

3. Generate 5000 Random Values Between 0 and 1

Using numpy's rand() function, we can quickly create 5000 random floats in the [0,1) range:

# Generate 5000 random values
random_values = np.random.rand(5000)

4. Find the Closest Larger Value in the sum Column

Since your sum column is sorted in ascending order, binary search is the most efficient way to find the first value larger than each random number. Here's how it works:

  • Extract the sum column into a list for faster lookup
  • Use bisect_right() to find the insertion point for each random value (this gives us the index of the first element in sum that's larger than the random value)
  • If a random value is bigger than all elements in sum, we'll default to the last row (since there's no larger value available)
# Extract sum values into a list for binary search
sum_values = original_df['sum'].tolist()

# Get indices of the closest larger values
selected_indices = []
for val in random_values:
    idx = bisect.bisect_right(sum_values, val)
    # Handle cases where no larger value exists
    if idx == len(sum_values):
        selected_indices.append(len(sum_values) - 1)
    else:
        selected_indices.append(idx)

5. Create the Final 5000-Row DataFrame

Now we'll use the indices we found to pull the corresponding rows from the original DataFrame, then reset the index to clean up the result:

# Generate the new DataFrame
new_df = original_df.iloc[selected_indices].reset_index(drop=True)

# Optional: Verify the result shape (should be 5000 rows, 2 columns)
print(new_df.shape)

Key Notes:

  • If your actual sum column isn't sorted, you'll need to sort it first (and keep track of original indices if you need to preserve the original row order)
  • The fallback to the last row for values larger than all sum entries is a practical default. If you want to handle these cases differently (e.g., skip them), you can adjust the loop logic accordingly

内容的提问来源于stack exchange,提问作者Сергей Гульдяшев

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:05:14