Scikit-learn监督学习:含非训练用字符串数据集的处理问询
处理带关联元数据的分类数据集
Got it, let's tackle this problem step by step. You've got a CSV with metadata (first name, last name, identifier) that you need to keep linked to each data entry, while only using the integer columns for your classifier training. Here's how to handle this cleanly:
方法1:使用NumPy(基于你提到的genfromtxt)
Since your CSV mixes string metadata and integer features, you'll need to define a mixed data type for genfromtxt to correctly parse all columns.
import numpy as np # 读取CSV,自动推断列类型,指定分隔符和编码 data = np.genfromtxt(dataInput, delimiter=',', dtype=None, encoding='utf-8') # 分离元数据和特征列 # 前3列是字符串元数据,后面的是整数特征 metadata = data[:, :3] # 格式为 (num_samples, 3),每个元素是(firstname, lastname, identifier) features = data[:, 3:].astype(np.int32) # 转换为整数数组,用于分类器训练 # 现在metadata和targetList是一一对应的,每个索引i对应同一个数据条目 # 比如查看第一个条目: print(f"Metadata: {metadata[0]}, Target: {targetList[0]}")
关键注意事项:
dtype=None让genfromtxt自动推断每列的类型,适配字符串和整数混合的场景- 添加
encoding='utf-8'避免Python新版本的字符串解码错误 - 必须保证
targetList和dataInput的行顺序完全一致,才能维持数据关联
方法2:使用Pandas(更直观的关联管理)
如果想更灵活地管理元数据、特征和目标的关联,Pandas DataFrames是理想选择——它把所有相关数据保存在一个结构里,避免手动对齐索引的麻烦。
import pandas as pd # 自动计算特征列数量 num_features = len(dataInput.split('\n')[0].split(',')) - 3 # 自定义列名 column_names = ['firstname', 'lastname', 'identifier'] + [f'feature_{i}' for i in range(num_features)] df = pd.read_csv(dataInput, names=column_names) # 直接把目标列加入DataFrame,自动完成关联 df['target'] = targetList # 提取用于训练的特征和目标 X = df.drop(['firstname', 'lastname', 'identifier', 'target'], axis=1) y = df['target'] # 轻松查看任意条目的完整信息 sample_idx = 0 print(f"Name: {df.loc[sample_idx, 'firstname']} {df.loc[sample_idx, 'lastname']}") print(f"ID: {df.loc[sample_idx, 'identifier']}") print(f"Features: {X.loc[sample_idx].values}") print(f"Target: {df.loc[sample_idx, 'target']}")
为什么选Pandas?
- 筛选、排序或修改数据时,元数据、特征和目标会自动保持关联
- 相比原生NumPy数组,更容易调试和检查数据集的细节
不管你选哪种方法,核心都是确保元数据、特征、目标列表的索引严格对应,这样后续训练完成后,你能轻松把分类结果和原始的名字/标识符关联起来。
内容的提问来源于stack exchange,提问作者Justin Crerar
相关产品推荐
相关产品推荐

