You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何向Pandas DataFrame高效批量新增多列?附图像像素列报错解决

Got it, let's break down your problem and fix this properly. First, that error you're seeing comes from trying to assign a 2D numpy array directly to a row slice in pandas—pandas expects the value shape to match exactly what you're assigning to, and your img_green_data is a (1, 6869376) array, while loc[example,2:] expects a 1D sequence. Plus, looping through every row to add columns is super inefficient for pandas—we can do this way faster with batch operations.

Step 1: Batch Extract All Green Channel Pixels

Instead of processing one image at a time and modifying the DataFrame row-by-row, we'll first load all images' green channel data into a single 2D numpy array (each row = one image's pixel vector). This is way faster and avoids shape mismatch issues.

import pandas as pd
import numpy as np
from PIL import Image

# Load original training metadata
df_train = pd.read_csv('../train.csv')
img_dir = "../input/train_images/"
total_pixels = 6869376  # Total pixels per image

# Helper function to extract flattened green channel from an image
def extract_green_pixels(img_id):
    img_path = f"{img_dir}{img_id}.png"
    with Image.open(img_path) as img:
        # Convert image to numpy array, grab green channel (index 1), flatten to 1D
        return np.array(img)[:, :, 1].flatten()

# Batch process all images into a 2D array (shape: [number of images, total_pixels])
pixel_data = np.array([extract_green_pixels(img_id) for img_id in df_train['id_code']])

Step 2: Merge Pixel Data with Original DataFrame

Next, we'll turn our pixel array into a DataFrame with meaningful column names, then concatenate it with the original df_train to get your desired df_train_new.

# Create column names for each pixel (e.g., pixel_0, pixel_1, ..., pixel_6869375)
pixel_columns = [f"pixel_{i}" for i in range(total_pixels)]

# Convert pixel array to a DataFrame
df_pixels = pd.DataFrame(pixel_data, columns=pixel_columns)

# Merge with original metadata to get the final DataFrame
df_train_new = pd.concat([df_train, df_pixels], axis=1)

Why This Works Better Than Your Original Code

  • No row-by-row modification: Pandas is optimized for batch operations—looping and using loc to assign columns will crawl to a halt if you have thousands of images.
  • Avoids shape errors: We're creating a properly structured 2D array first, so when we convert it to a DataFrame, it aligns perfectly with the original data during concatenation.
  • Cleaner, maintainable code: The logic is split into clear steps, making it easier to debug or modify later.

Bonus: Speed Up Loading for Large Datasets

If you have a huge number of images, use multiprocessing to load them in parallel (utilizing all your CPU cores):

from multiprocessing import Pool

# Use multiprocessing to process images in parallel
with Pool() as pool:
    pixel_list = pool.map(extract_green_pixels, df_train['id_code'].tolist())
pixel_data = np.array(pixel_list)

Answering Your Second Question: Batch Adding Columns to Pandas

The simplest ways to add multiple columns at once are:

  1. Concatenate with a new DataFrame: As we did above—create a DataFrame with your new columns, then use pd.concat([original_df, new_df], axis=1). Perfect when your new data is a 2D array aligned with the original rows.
  2. Direct array assignment: If you already have column names defined, you can assign the array directly:
    df_train[pixel_columns] = pixel_data
    
    Pandas will automatically map each column in the array to the corresponding name in your list.
  3. df.assign() for computed columns: For columns derived from existing data, use this clean syntax:
    df_train = df_train.assign(
        doubled_diagnosis=lambda x: x['diagnosis'] * 2,
        diagnosis_plus_5=lambda x: x['diagnosis'] + 5
    )
    

内容的提问来源于stack exchange,提问作者Economister

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:41:24