Python DataFrame:多列与指定列相乘并保存至新列的高效实现
Efficiently Multiply Columns B-Z with Column A and Add Results as R1-R25
Great question! When working with larger datasets like your 160-column × 16k-row DataFrame, efficiency hinges on using Pandas' vectorized operations (implemented in C, far faster than Python loops) and avoiding unnecessary memory overhead. Here's a streamlined, high-performance solution:
Step-by-Step Implementation
- Target the columns to multiply: First, get all columns except
A(B through Z, or the 159 columns after A in your full dataset). Dynamic selection ensures your code works even if column order changes. - Compute the vectorized product: Use Pandas'
mul()method (equivalent tomultiply()) to broadcast columnAacross all target columns—this is way faster than looping through columns manually. - Rename the result columns: Assign the new column names
R1throughR25(orR1throughR159for your full dataset). - Add results to the original DataFrame: Use direct assignment to avoid creating an extra DataFrame copy, which saves memory.
import pandas as pd # Get all columns except 'A' (adjust if 'A' isn't the first column) cols_to_multiply = df.columns[df.columns != 'A'] # Calculate product of each target column with 'A' (vectorized operation) product_columns = df[cols_to_multiply].mul(df['A'], axis=0) # Rename columns to R1, R2, ..., Rn product_columns.columns = [f'R{i+1}' for i in range(len(product_columns.columns))] # Add the new columns to the original DataFrame (memory-efficient assignment) df[product_columns.columns] = product_columns
Why This Is Better Than Your Original Approach
- No extra DataFrame copies: Directly assigning the new columns to
dfavoids the overhead of creating a separate result DataFrame and then concatenating it (which would duplicate data temporarily). - Full vectorization: The
mul()operation runs entirely in optimized C code, so it’s orders of magnitude faster than any Python-level loop over columns. - Scalability: This code works seamlessly for your 160-column dataset (it will generate
R1toR159automatically) without any modifications.
Alternative (If You Prefer concat)
If you’d rather create a new DataFrame instead of modifying the original, pd.concat is still efficient (though slightly less memory-friendly than direct assignment):
df_with_results = pd.concat([df, product_columns], axis=1)
内容的提问来源于stack exchange,提问作者Hugo Ribeiro
相关产品推荐
相关产品推荐

