如何将文本分类模型的Numpy数组输出保存为含X_test与预测Y_test的CSV文件
Hey there! Let's walk through how to save your text classification results—including your X_test data and the predicted labels (Y_pred, which you referred to as Y_test)—into a clean CSV file. Since you've already got your NumPy arrays set up and used a Pipeline for classification, we just need to bridge those arrays to a CSV-friendly format.
Step 1: Ensure You Have Your Predictions Ready
First, make sure you've generated your predicted labels using your trained classifier. For example:
# Assuming your trained classifier is named `clf` y_pred = clf.predict(X_test)
Step 2: Combine X_test and Predictions into a DataFrame
Pandas makes it super easy to merge your data and save it. Here are two common scenarios:
Scenario 1: X_test is Raw Text Data (from a CSV)
If your X_test comes from a raw test CSV (with original text columns you want to preserve), load that test data first, then add the predictions as a new column:
import pandas as pd # Load your original test data (adjust the file path and encoding as needed) test_data = pd.read_csv('your_test_data.csv', encoding='latin1', dtype={'SourcePath': str}) # Add the predicted labels as a new column test_data['Y_pred'] = y_pred # Save to CSV (index=False skips the auto-generated row numbers) test_data.to_csv('classification_results.csv', index=False, encoding='latin1')
Scenario 2: X_test is a NumPy Feature Array
If X_test is a NumPy array of extracted features (e.g., TF-IDF vectors), convert it to a DataFrame first, then merge with predictions:
import pandas as pd import numpy as np # Convert NumPy array to DataFrame (add column names if you know your feature names) x_test_df = pd.DataFrame(X_test, columns=[f'feature_{i}' for i in range(X_test.shape[1])]) # Add the predicted labels column x_test_df['Y_pred'] = y_pred # Save to CSV x_test_df.to_csv('feature_results_with_predictions.csv', index=False)
Step 3: Save NumPy Arrays Directly to CSV (If Needed)
If you just need to save the raw NumPy arrays (without combining them into a full dataset), you can use either NumPy's built-in function or pandas:
- Save
X_testarray:# Using NumPy np.savetxt('x_test_array.csv', X_test, delimiter=',', fmt='%f') # Use %s for string features # Using pandas (better for readability with column names) pd.DataFrame(X_test).to_csv('x_test_array.csv', index=False) - Save
y_predarray:# Using NumPy np.savetxt('y_pred_array.csv', y_pred, delimiter=',', fmt='%s') # Use %s for categorical labels # Using pandas pd.DataFrame(y_pred, columns=['Y_pred']).to_csv('y_pred_array.csv', index=False)
Quick Tips
- Encoding: Stick with
encoding='latin1'orutf-8to avoid character encoding issues with text data. - Data Types: Use
dtypeparameters when loading CSVs to preserve data types (like keepingSourcePathas a string, as you did in your code). - Clean Output: Always use
index=Falsewhen saving to CSV unless you specifically need the row index.
内容的提问来源于stack exchange,提问作者PiotrK

