如何向Pandas DataFrame高效批量新增多列?附图像像素列报错解决
Got it, let's break down your problem and fix this properly. First, that error you're seeing comes from trying to assign a 2D numpy array directly to a row slice in pandas—pandas expects the value shape to match exactly what you're assigning to, and your img_green_data is a (1, 6869376) array, while loc[example,2:] expects a 1D sequence. Plus, looping through every row to add columns is super inefficient for pandas—we can do this way faster with batch operations.
Step 1: Batch Extract All Green Channel Pixels
Instead of processing one image at a time and modifying the DataFrame row-by-row, we'll first load all images' green channel data into a single 2D numpy array (each row = one image's pixel vector). This is way faster and avoids shape mismatch issues.
import pandas as pd import numpy as np from PIL import Image # Load original training metadata df_train = pd.read_csv('../train.csv') img_dir = "../input/train_images/" total_pixels = 6869376 # Total pixels per image # Helper function to extract flattened green channel from an image def extract_green_pixels(img_id): img_path = f"{img_dir}{img_id}.png" with Image.open(img_path) as img: # Convert image to numpy array, grab green channel (index 1), flatten to 1D return np.array(img)[:, :, 1].flatten() # Batch process all images into a 2D array (shape: [number of images, total_pixels]) pixel_data = np.array([extract_green_pixels(img_id) for img_id in df_train['id_code']])
Step 2: Merge Pixel Data with Original DataFrame
Next, we'll turn our pixel array into a DataFrame with meaningful column names, then concatenate it with the original df_train to get your desired df_train_new.
# Create column names for each pixel (e.g., pixel_0, pixel_1, ..., pixel_6869375) pixel_columns = [f"pixel_{i}" for i in range(total_pixels)] # Convert pixel array to a DataFrame df_pixels = pd.DataFrame(pixel_data, columns=pixel_columns) # Merge with original metadata to get the final DataFrame df_train_new = pd.concat([df_train, df_pixels], axis=1)
Why This Works Better Than Your Original Code
- No row-by-row modification: Pandas is optimized for batch operations—looping and using
locto assign columns will crawl to a halt if you have thousands of images. - Avoids shape errors: We're creating a properly structured 2D array first, so when we convert it to a DataFrame, it aligns perfectly with the original data during concatenation.
- Cleaner, maintainable code: The logic is split into clear steps, making it easier to debug or modify later.
Bonus: Speed Up Loading for Large Datasets
If you have a huge number of images, use multiprocessing to load them in parallel (utilizing all your CPU cores):
from multiprocessing import Pool # Use multiprocessing to process images in parallel with Pool() as pool: pixel_list = pool.map(extract_green_pixels, df_train['id_code'].tolist()) pixel_data = np.array(pixel_list)
Answering Your Second Question: Batch Adding Columns to Pandas
The simplest ways to add multiple columns at once are:
- Concatenate with a new DataFrame: As we did above—create a DataFrame with your new columns, then use
pd.concat([original_df, new_df], axis=1). Perfect when your new data is a 2D array aligned with the original rows. - Direct array assignment: If you already have column names defined, you can assign the array directly:
Pandas will automatically map each column in the array to the corresponding name in your list.df_train[pixel_columns] = pixel_data df.assign()for computed columns: For columns derived from existing data, use this clean syntax:df_train = df_train.assign( doubled_diagnosis=lambda x: x['diagnosis'] * 2, diagnosis_plus_5=lambda x: x['diagnosis'] + 5 )
内容的提问来源于stack exchange,提问作者Economister

