You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在循环中解析鸢尾花txt文件创建指定列的Numpy数组?

Fixing Your Iris Data Numpy Array Creation Issue

Hey there! Let's work through why your current code isn't working and fix it up properly.

First, let's break down the problems in your original code:

  • You're overwriting feature_table every loop: Instead of accumulating rows, you create a new array each time, so you'll only ever end up with the last line's data.
  • Incorrect numpy array initialization: np.array([currentline[0]],[currentline[3]],[currentline[4]]) is not valid syntax—numpy expects a single nested structure for rows/columns, not separate lists.
  • No type conversion: The values from split(",") are strings, but your first two features should be numerical (float/int) for any future analysis.
  • Unsafe file handling: You're opening the file but not explicitly closing it—using a with statement is cleaner and safer.
  • Inconsistent module naming: You imported numpy as np but tried to use numpy.array in commented lines, which would throw a NameError.

Here's the corrected code using np.append as you requested:

import numpy as np

# Initialize an empty 2D array to hold our data (dtype=object lets us mix numbers and strings)
feature_table = np.empty((0, 3), dtype=object)

# Use 'with' to safely open/close the file
with open("iris_data.txt", "r") as iris_data:
    for line in iris_data:
        # Strip newline characters and split the line by commas
        currentline = line.strip().split(",")
        
        # Skip any empty lines in the file
        if not currentline:
            continue
            
        # Convert numerical features to float, clean up the category name
        sepal_length = float(currentline[0])
        petal_width = float(currentline[3])
        iris_name = currentline[4].strip()  # Remove any leftover whitespace/newlines
        
        # Create a 2D row array (matches the shape of feature_table for appending)
        current_row = np.array([[sepal_length, petal_width, iris_name]])
        
        # Append the row to our feature table along the row axis (axis=0)
        feature_table = np.append(feature_table, current_row, axis=0)

# Check the result
print(feature_table)
print(feature_table.shape)

A more efficient alternative (avoiding looped np.append):

Looping with np.append can be slow for large datasets because numpy reallocates memory every time you append. A better approach is to collect all rows in a Python list first, then convert to a numpy array once:

import numpy as np

# Start with an empty Python list to collect rows
data_rows = []

with open("iris_data.txt", "r") as iris_data:
    for line in iris_data:
        currentline = line.strip().split(",")
        if not currentline:
            continue
            
        sepal_length = float(currentline[0])
        petal_width = float(currentline[3])
        iris_name = currentline[4].strip()
        
        data_rows.append([sepal_length, petal_width, iris_name])

# Convert the list to a numpy array once
feature_table = np.array(data_rows, dtype=object)

print(feature_table)
print(feature_table.shape)

This will give you the structured array you need, with the first two columns as numerical features and the third as the iris category name.

内容的提问来源于stack exchange,提问作者Emre Unsal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:50:21