如何基于DataFrame现有列批量生成计算后的新列?
Hey there! Let's get this sorted for you. Your goal is to generate weighted columns for every date column in your sega_df using the formula (100 - [5*eachcell])*0.2, with new columns named Weighted_<original-col-name>, and optionally cap negative values at 0. Here's how to do it efficiently:
First, let's fix your original approach issues
Your current loop has a couple of problems:
- You're iterating over row indices instead of columns, which isn't what you need
- You keep overwriting a single
Weightedcolumn instead of creating new ones for each original column - The line
sega_df[sega_df < 0] = 0modifies all columns (including your original data) instead of just the new weighted columns
Solution 1: Simple column-wise loop (easy to read)
This approach loops through each column, calculates the weighted values, and adds the new column directly to your DataFrame:
# Loop through each column in your DataFrame for col in sega_df.columns: # Create the new column name weighted_col_name = f"Weighted_{col}" # Apply your formula to the column sega_df[weighted_col_name] = (100 - 5 * sega_df[col]) * 0.2 # Optional: Cap negative values at 0 (remove this line if you want to keep negatives) sega_df[weighted_col_name] = sega_df[weighted_col_name].clip(lower=0)
Solution 2: Vectorized operation (faster for large datasets)
If you're working with a big DataFrame, this vectorized method is more efficient—it calculates all weighted columns at once without explicit loops:
import pandas as pd # Calculate weighted values for all columns in one go weighted_data = (100 - 5 * sega_df) * 0.2 # Rename columns to add the Weighted_ prefix weighted_data.columns = [f"Weighted_{col}" for col in weighted_data.columns] # Optional: Cap negatives at 0 (remove if not needed) weighted_data = weighted_data.clip(lower=0) # Merge the original DataFrame with the new weighted columns sega_df = pd.concat([sega_df, weighted_data], axis=1)
Example Output
Using your sample data:
# First, recreate your original DataFrame data = { "2016-11-01": [12.0, 5.0], "2016-12-01": [3.0, 23.0] } sega_df = pd.DataFrame(data, index=["Sonic", "Shadow"]) sega_df.index.name = "Character"
After running either solution (without the clip step to keep negatives), you'll get exactly your desired output:
2016-11-01 2016-12-01 Weighted_2016-11-01 Weighted_2016-12-01 Character Sonic 12.0 3.0 8.0 17.0 Shadow 5.0 23.0 15.0 -3.0
内容的提问来源于stack exchange,提问作者florence-y

