numpy genfromtxt读取混合类型数据如何返回二维对象数组
NumPy读取混合类型文本生成object二维数组问题
问题场景
需要从列数不固定的文本文件中导入数据,存储为数组嵌套数组的格式:
- 数据第一列始终为字符串类型
- 后续三列为整数类型
使用np.genfromtxt读取时,设置dtype=(object,int,int,int)参数仅能得到元组组成的结构化数组,无法得到目标二维数组。
测试代码
from io import StringIO import numpy as np new_string = StringIO("01/23/2020, 32, 0, 2 \n01/31/2020' ,436 ,0 ,10") new_result = np.genfromtxt(new_string, dtype=(object,int,int,int), encoding="unicode" , delimiter=",") print("File data:",new_result )
当前输出
File data: [('01/23/2020', 32, 0, 2) ("01/31/2020' ", 436, 0, 10)]
期望输出
需要得到dtype为object的二维数组,格式如下:
[['01/23/2020' 32 0 2] ['01/31/2020' 436 0 10]]
要求读取结果与如下手动构造的数组做相等判断时返回True:
new_result == np.array( [['01/23/2020',32,0,2], ['01/31/2020', 436, 0, 10]],dtype=object)
原因说明
给np.genfromtxt传入多类型组成的dtype元组时,NumPy会默认生成结构化数组,数组每个元素是存储不同类型值的元组,而非通用object类型的二维数组,这是当前输出不符合预期的核心原因。
解决方法
- 方法1:读取结构化数组后直接转换
适合列数固定的场景,代码改动最小:from io import StringIO import numpy as np new_string = StringIO("01/23/2020, 32, 0, 2 \n01/31/2020' ,436 ,0 ,10") new_result = np.genfromtxt(new_string, dtype=(object,int,int,int), encoding="unicode", delimiter=",") # 转为列表后重新生成object类型二维数组 new_result = np.array(new_result.tolist(), dtype=object) - 方法2:分字段读取后拼接
适配列数不固定的场景,扩展性更好:from io import StringIO import numpy as np new_string = StringIO("01/23/2020, 32, 0, 2 \n01/31/2020' ,436 ,0 ,10") # 单独读取第一列字符串 col_str = np.genfromtxt(new_string, delimiter=",", usecols=0, dtype=object, encoding="unicode") new_string.seek(0) # 读取后续所有整数列,列数变化时只需要调整usecols范围即可 col_int = np.genfromtxt(new_string, delimiter=",", usecols=range(1, 4), dtype=int) # 拼接为目标二维数组 new_result = np.column_stack([col_str, col_int]).astype(object)
两种方法得到的new_result和手动构造的object数组做相等判断时,所有元素都会返回True,符合预期。
内容的提问来源于stack exchange,提问作者Ohad Sharet
相关产品推荐
相关产品推荐

