Pandas:如何高效地在DataFrame每行后重复插入指定行
Efficiently Insert Repeated Rows After Each Row in a Pandas DataFrame
Great question—iterating row-by-row with iterrows() and repeatedly concatenating small DataFrames is terrible for performance with large datasets. Let's use pandas and numpy's vectorized operations to solve this way faster, no loops required:
Optimized Solution
import pandas as pd import numpy as np # Your original DataFrame and row to insert df = pd.DataFrame({ 'index': [0, 1, 2], 'val1': [0.2, 0.5, 0.3], 'val2': [0.1, 0.7, 0.4] }) myrow = pd.DataFrame({ 'index': [-1], 'val1': [0.9], 'val2': [0.9] }) # Step 1: Create a base DataFrame where each original row is repeated 4 times result = df.loc[df.index.repeat(4)].reset_index(drop=True) # Step 2: Create a mask for positions that need to be replaced with myrow # (every position except the first in each group of 4 rows) mask = np.arange(len(result)) % 4 != 0 # Step 3: Batch-replace the masked positions with myrow's values result.loc[mask] = myrow.values[0]
How This Works
- Repeat Rows Efficiently:
df.index.repeat(4)generates an index that repeats each original row's index 4 times. When we index the DataFrame with this, we get a structure where each original row is duplicated 4 times—this is done in C-level code, way faster than Python loops. - Target Replacement Positions: The
maskuses numpy's vectorized arithmetic to mark every position that isn't the first row in each 4-row block (these are the spots where we want to insertmyrow). - Batch Update: Instead of modifying rows one by one, we use
result.loc[mask] = myrow.values[0]to overwrite all target positions in a single vectorized operation. This avoids the overhead of repeatedconcatcalls and row-level operations.
Performance Comparison
For a DataFrame with 100,000 rows, this method will run in milliseconds, whereas the loop-based approach would take minutes. Vectorized operations are the key to pandas performance—always prioritize them over row-wise loops when possible.
Note
If myrow is a pandas Series instead of a single-row DataFrame, you can simplify the last line to result.loc[mask] = myrow.values (no need for [0]).
内容的提问来源于stack exchange,提问作者Student
相关产品推荐
相关产品推荐

