Python中Numpy数组追加/拼接异常:空数组问题求助
解决Numpy数组为空的问题
Hey there, let's figure out why your numpy array is coming up empty and fix this issue step by step. Looking at your code snippet, there are a few key mistakes that are causing this problem:
问题分析
- 文件读取未获取内容:You just opened the file with
open("auto"+x)but never actually read the text inside it.CountVectorizer.fit_transform()needs actual text data, not just a file handle. - 空数组未被填充:You initialized
list1=np.array([])but never added any processed data to it. Plus, your code got truncated atX=np.ar..., so even if you were trying to convert data to a numpy array, it wasn't completed. - 稀疏矩阵未转换:
fit_transform()returns a sparsescipymatrix, not a numpy array directly. You need to convert it explicitly. - 宽泛的异常处理:Your
tryblock doesn't catch specific errors, so you might be skipping files without knowing why (like missing files).
修复后的代码
Here's a revised version of your code that fixes all these issues:
from sklearn.feature_extraction.text import CountVectorizer from sklearn.naive_bayes import GaussianNB import numpy as np # Use a list to store each file's feature array (more efficient than numpy concatenation in loops) data_collection = [] for x in range(101563, 103807): file_path = f"auto{x}" try: # Safely read the file content using a with statement (auto-closes the file) with open(file_path, 'r', encoding='utf-8') as file: text_data = file.read() # Initialize CountVectorizer and process the text count_vectorizer = CountVectorizer() # Pass the text as a list (fit_transform expects an iterable of documents) sparse_features = count_vectorizer.fit_transform([text_data]) # Convert sparse matrix to a dense numpy array feature_array = sparse_features.toarray() # Add the processed array to our collection data_collection.append(feature_array) except FileNotFoundError: print(f"Warning: File {file_path} doesn't exist, skipping.") except Exception as e: print(f"Error processing {file_path}: {str(e)}") # Combine all individual arrays into one large numpy array final_numpy_array = np.vstack(data_collection) # Verify the result print(f"Final array shape: {final_numpy_array.shape}") print("First 3 samples:") print(final_numpy_array[:3])
额外优化建议
If you want all files to use the same vocabulary (so every feature array has the same dimensions), you should fit the CountVectorizer once on all text data instead of per file. This avoids shape mismatches when combining arrays:
# First, collect all text data from valid files all_texts = [] valid_files = [] for x in range(101563, 103807): file_path = f"auto{x}" try: with open(file_path, 'r', encoding='utf-8') as file: all_texts.append(file.read()) valid_files.append(file_path) except Exception as e: print(f"Skipping {file_path}: {str(e)}") # Fit the vectorizer on all texts to create a shared vocabulary count_vectorizer = CountVectorizer() sparse_features = count_vectorizer.fit_transform(all_texts) # Convert to a single numpy array final_numpy_array = sparse_features.toarray() print(f"Final array shape (shared vocabulary): {final_numpy_array.shape}")
This ensures every entry in your numpy array has the same number of features, which is crucial if you're planning to use this data with models like GaussianNB.
内容的提问来源于stack exchange,提问作者hhc
相关产品推荐
相关产品推荐

