You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取含28×28图像数据的CSV文件:解决numpy读取失败并转784数组

How to Load and Convert Your CSV Image Data to 784-Dimensional Arrays

Hey there! Let's work through this problem together. The reason numpy.load and numpy.loadtxt are failing is almost certainly because your first column isn't a single numerical value—it's a string packed with 784 integers (the 28×28 pixel data) all in one cell. Numpy's default loaders can't parse that nested structure automatically, so we need to handle the string splitting manually. Here are two reliable approaches:

Approach 1: Use Python's Built-in csv Module

This is great if you want to stick to standard libraries without extra dependencies:

import numpy as np
import csv

# Initialize a list to hold all processed image arrays
image_dataset = []

# Open your CSV file (replace 'your_data.csv' with your actual file path)
with open('your_data.csv', 'r') as csv_file:
    csv_reader = csv.reader(csv_file)
    
    # Skip the header row if your CSV has one (remove this line if there's no header)
    next(csv_reader)
    
    for row in csv_reader:
        # Split the first column's string into individual pixel strings
        # Adjust the delimiter (',' here) if your pixels are separated by spaces/semicolons
        pixel_strings = row[0].split(',')
        # Convert the strings to integers and turn into a numpy array
        pixel_array = np.array(pixel_strings, dtype=np.int32)
        
        # Quick sanity check to ensure we have exactly 784 pixels
        if pixel_array.shape != (784,):
            print(f"Warning: Row has {len(pixel_array)} pixels instead of 784")
            continue
        
        image_dataset.append(pixel_array)

# Convert the list of arrays into a single 2D numpy array (shape: [number_of_samples, 784])
final_image_array = np.array(image_dataset)

Approach 2: Use Pandas (Simpler and Faster for Large Datasets)

If you don't mind using pandas (a common data handling library), this method is more concise:

import numpy as np
import pandas as pd

# Load the CSV file. Adjust header=None and names if your file has no column headers
df = pd.read_csv('your_data.csv')  # For no headers: pd.read_csv('your_data.csv', header=None, names=['image', 'label'])

# Define a helper function to process each image string
def parse_image(s):
    # Again, adjust the delimiter to match your CSV's format
    return np.array(s.split(','), dtype=np.int32)

# Apply the function to the first column and convert to a numpy array
image_array = np.array(df.iloc[:, 0].apply(parse_image).tolist())

# Verify the shape should be (number_of_samples, 784)
print(f"Processed image array shape: {image_array.shape}")

Key Notes:

  • Delimiter Check: If your pixel data is separated by spaces instead of commas, replace split(',') with split() (which handles any whitespace). If it's another character (like ;), update the split argument accordingly.
  • Data Types: I used np.int32 since your data is integers, but you can adjust to np.uint8 if your pixel values are 0-255 (common for images) to save memory.

内容的提问来源于stack exchange,提问作者Code Developer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:07:09