You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于Pandas DataFrame列值生成指定数量的均匀分布新行?

Solution to Generate Evenly Spaced Rows from DataFrame Entries

Got it, let's break down how to solve this problem efficiently—even for large DataFrames with extra, unrelated columns. The core idea is to generate a sequence of evenly spaced values for each row, then expand those sequences into individual rows while preserving any other columns you need to keep.

Step-by-Step Implementation

First, let's start with your sample input to verify the solution works as expected:

import pandas as pd

# Sample input DataFrame
df = pd.DataFrame({'base_number': [2], 'std_dev': [1]})
df['amount_needed'] = 5
df['upper_bound'] = df['base_number'] + df['std_dev']
df['lower_bound'] = df['base_number'] - df['std_dev']

1. Generate Evenly Spaced Sequences

For each row, we use numpy.linspace to create a list of values between lower_bound and upper_bound, with exactly amount_needed elements. We'll add this as a new column to the original DataFrame:

# Create a column containing the evenly spaced values list
df['new_base_values'] = df.apply(
    lambda row: pd.np.linspace(
        start=row['lower_bound'],
        stop=row['upper_bound'],
        num=row['amount_needed']
    ),
    axis=1
)

2. Expand Lists into Individual Rows

Next, we use explode() to turn each element in the list into a separate row. This will automatically carry over all other columns from the original row to each new row—perfect for retaining metadata from large datasets:

# Explode the list column into distinct rows
df_expanded = df.explode('new_base_values').reset_index(drop=True)

3. Clean Up the Final Output

Finally, we can rename the column back to base_number and drop any helper columns (like lower_bound, upper_bound, etc.) if they're not needed in your final result:

# Rename and remove unnecessary columns
df_new = df_expanded.rename(columns={'new_base_values': 'base_number'})
df_new = df_new.drop(columns=['std_dev', 'amount_needed', 'lower_bound', 'upper_bound'])

Sample Output

Running this code on your sample input will produce exactly the expected result:

base_number
0          1.0
1          1.5
2          2.0
3          2.5
4          3.0

Handling Large DataFrames & Extra Columns

This approach works seamlessly for large datasets with additional columns. For example, if you have a DataFrame with multiple rows and irrelevant metadata:

# Example with multiple rows and extra columns
df_large = pd.DataFrame({
    'base_number': [2, 4],
    'std_dev': [1, 2],
    'amount_needed': [5, 3],
    'category': ['X', 'Y'],
    'record_id': [101, 102]
})
df_large['upper_bound'] = df_large['base_number'] + df_large['std_dev']
df_large['lower_bound'] = df_large['base_number'] - df_large['std_dev']

Applying the same steps will generate 5 rows for the first entry and 3 for the second, with all extra columns (category, record_id) preserved for each new row.

Key Notes

  • numpy.linspace ensures values are evenly distributed, including both the lower and upper bounds (matching your example exactly).
  • explode() is far more efficient than looping through each row manually, which is critical for large datasets.
  • All original columns are retained unless explicitly dropped, so you don’t lose any context from the original rows.

内容的提问来源于stack exchange,提问作者Mars

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 16:27:41