如何使用Pandas将Series追加至DataFrame?及优化DataFrame行迭代与Series重构代码效率方案
Hey there! Let's break down your two Pandas questions with practical, efficient solutions.
There are a few straightforward ways to do this, depending on whether you're adding the Series as a new column or a new row:
追加为新列
If your Series has a matching index with the DataFrame, you can directly assign it, or use pd.concat:
import pandas as pd # 创建示例DataFrame和Series df = pd.DataFrame({'A': [1, 2, 3], 'B': [4, 5, 6]}) new_col = pd.Series([7, 8, 9], name='C') # 方法1:直接赋值(最简单) df['C'] = new_col # 方法2:用pd.concat df = pd.concat([df, new_col], axis=1)
追加为新行
For adding a row, make sure the Series index matches the DataFrame's column names. You can use pd.concat (by converting the Series to a single-row DataFrame) or df.loc:
new_row = pd.Series([10, 11, 12], index=['A', 'B', 'C']) # 方法1:pd.concat(适合批量追加) df = pd.concat([df, new_row.to_frame().T], axis=0, ignore_index=True) # 方法2:df.loc(适合单行追加) df.loc[len(df)] = new_row
⚠️ Note: The old df.append() method is deprecated in Pandas 2.0+, so stick to the methods above.
Your current code is slow for two key reasons:
iterrows()is a slow, row-by-row iterator that doesn't leverage Pandas' vectorized capabilities- Calling
pd.concat()inside the loop creates a new DataFrame every iteration, which is extremely inefficient for 4k+ rows
Here are two much faster solutions:
方案1:全向量化操作(推荐,最快)
Pandas is designed for bulk operations on entire DataFrames. We can rewrite your logic without any loops:
df_prices = get_prices_df() # Step 1: Calculate rank for every row in df_prices (all columns at once) rank_df = df_prices.rank(axis=1, ascending=False, na_option='bottom') # Step 2: Keep only values where df_mask is True, then check if rank ≤10 container = (rank_df <= 10) & df_mask # Step 3: Fill any remaining NaNs with False (though the above logic should already handle this) container = container.fillna(False)
This approach processes the entire dataset in bulk, which will cut your runtime from minutes to seconds.
方案2:优化循环(如果必须保留逐行逻辑)
If you need to keep some custom per-row logic, avoid repeated concat by collecting results in a list first, then merging once:
df_prices = get_prices_df() container_list = [] for idx, row in df_mask.iterrows(): cols = row[row == True].index series = df_prices.loc[idx, cols].rank(axis=1, ascending=False, na_option='bottom').le(10) container_list.append(series) # Merge all collected Series into one DataFrame at once container = pd.concat(container_list, axis=0).unstack().fillna(False)
This reduces the number of expensive DataFrame copies drastically, making the loop way faster than your original code.
内容的提问来源于stack exchange,提问作者Florent

