You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:如何高效地在DataFrame每行后重复插入指定行

Efficiently Insert Repeated Rows After Each Row in a Pandas DataFrame

Great question—iterating row-by-row with iterrows() and repeatedly concatenating small DataFrames is terrible for performance with large datasets. Let's use pandas and numpy's vectorized operations to solve this way faster, no loops required:

Optimized Solution

import pandas as pd
import numpy as np

# Your original DataFrame and row to insert
df = pd.DataFrame({
    'index': [0, 1, 2],
    'val1': [0.2, 0.5, 0.3],
    'val2': [0.1, 0.7, 0.4]
})
myrow = pd.DataFrame({
    'index': [-1],
    'val1': [0.9],
    'val2': [0.9]
})

# Step 1: Create a base DataFrame where each original row is repeated 4 times
result = df.loc[df.index.repeat(4)].reset_index(drop=True)

# Step 2: Create a mask for positions that need to be replaced with myrow
# (every position except the first in each group of 4 rows)
mask = np.arange(len(result)) % 4 != 0

# Step 3: Batch-replace the masked positions with myrow's values
result.loc[mask] = myrow.values[0]

How This Works

  1. Repeat Rows Efficiently: df.index.repeat(4) generates an index that repeats each original row's index 4 times. When we index the DataFrame with this, we get a structure where each original row is duplicated 4 times—this is done in C-level code, way faster than Python loops.
  2. Target Replacement Positions: The mask uses numpy's vectorized arithmetic to mark every position that isn't the first row in each 4-row block (these are the spots where we want to insert myrow).
  3. Batch Update: Instead of modifying rows one by one, we use result.loc[mask] = myrow.values[0] to overwrite all target positions in a single vectorized operation. This avoids the overhead of repeated concat calls and row-level operations.

Performance Comparison

For a DataFrame with 100,000 rows, this method will run in milliseconds, whereas the loop-based approach would take minutes. Vectorized operations are the key to pandas performance—always prioritize them over row-wise loops when possible.

Note

If myrow is a pandas Series instead of a single-row DataFrame, you can simplify the last line to result.loc[mask] = myrow.values (no need for [0]).

内容的提问来源于stack exchange,提问作者Student

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 23:07:39