You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于指定列重复DataFrame行的Python实现求助

How to Repeat Rows in a Pandas DataFrame Correctly

It sounds like you’re trying to duplicate rows in your DataFrame but hitting roadblocks with your current code. Let’s break down the issues with your attempts and share working solutions for both common use cases: repeating every row a fixed number of times, or repeating rows based on a value in a column (like your mas_id).


Scenario 1: Repeat every row a fixed number of times

If you want each row in your DataFrame to repeat the same number of times (say, final_n times), here are solid fixes for your approaches:

Working Method 1: Using index.repeat

This is the cleanest approach and should work if final_n is an integer:

import pandas as pd

# Sample DataFrame
df = pd.DataFrame({'col1': [1, 2, 3], 'col2': ['a', 'b', 'c']})
final_n = 3  # Number of times to repeat each row

# Repeat each row final_n times
new_df = df.loc[df.index.repeat(final_n)].reset_index(drop=True)
print(new_df)

Your original code here might have failed if final_n wasn’t properly defined, or if you accidentally referenced a column instead of a fixed integer.

Working Method 2: Fixing the concat approach

Your concat attempt repeats the entire DataFrame final_n times (stacking the whole DF on top of itself), not each row individually. If that’s actually what you want, it works—but if you need per-row repetition, stick with the index.repeat method above.

Working Method 3: Fixing the numpy approach

Your numpy method works but loses column names and data types. To retain those:

import numpy as np

new_df = pd.DataFrame(np.repeat(df.values, final_n, axis=0), columns=df.columns)
# Restore original data types
new_df = new_df.astype(df.dtypes)

Scenario 2: Repeat rows based on a column value (e.g., mas_id)

If you want each row to repeat a number of times specified by the mas_id column (like row 1 repeats 2 times, row 2 repeats 5 times, etc.), your first attempt was on the right track—but let’s ensure it’s set up correctly:

Working Code

# Sample DataFrame with mas_id column (number of repetitions per row)
df = pd.DataFrame({'col1': [1, 2, 3], 'mas_id': [2, 3, 1]})

# Repeat each row according to mas_id
new_df = df.loc[df.index.repeat(df['mas_id'])].reset_index(drop=True)
print(new_df)

Possible reasons this might have failed for you:

  • The mas_id column contains non-integer values (e.g., floats or strings). Convert it to integers first with df['mas_id'] = df['mas_id'].astype(int)
  • The mas_id column has missing values (NaNs). Drop or fill those before using repeat: df['mas_id'] = df['mas_id'].fillna(0).astype(int)

All these methods should work depending on your specific use case. Let me know if you need further adjustments based on your exact DataFrame structure!

内容的提问来源于stack exchange,提问作者user19856161

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 13:39:06