如何用np.genfromtxt读取含hh:mm:ss格式时间列的文本文件?
问题描述
有如下格式的文本文件:
index timestamp polarisation current (A) signal (V) head temperature (°C) head relat.humidity (%RH) MUGS temperature (°C) laser voltage (V) laser current (A) driver temperature (°C) 0 16:11:24 0.4 0.0006019 26.51 43.5 32.0 11.37 0.3922 26.5 1 16:11:29 0.402 0.0006286 26.51 43.5 32.5 11.41 0.3972 31.5 2 16:11:34 0.404 0.0005828 26.51 43.5 32.5 11.42 0.4048 32.5 3 16:11:38 0.406 0.0006139 26.51 43.5 32.5 11.39 0.3984 32.5
使用以下代码读取文件时,第二列(timestamp)全为NaN:
with open(universal_path,'rt'): values = np.genfromtxt( universal_path, delimiter="", skip_header = 1, encoding='unicode_escape')
尝试设置dtype=None仅能读取第一列;设置dtype = [int, str, float, float, float, float, float, float, float, float]同样仅能读取第一列,需要实现将第二列以字符串形式正确读取。
解决方案
方法1:使用结构化数据类型(推荐)
numpy需要通过结构化数据类型明确指定每一列的类型,而非简单的类型列表。示例代码:
import numpy as np # 定义对应每一列的结构化数据类型 data_type = np.dtype([ ('index', int), ('timestamp', 'U8'), # U8表示长度为8的Unicode字符串,适配"HH:MM:SS"格式 ('polarisation_current', float), ('signal', float), ('head_temperature', float), ('head_humidity', float), ('mugs_temperature', float), ('laser_voltage', float), ('laser_current', float), ('driver_temperature', float) ]) # 读取文件,delimiter=None让numpy自动识别任意空白分隔符 values = np.genfromtxt( universal_path, dtype=data_type, skip_header=1, encoding='unicode_escape', delimiter=None ) # 访问第二列的时间数据 print(values['timestamp'])
方法2:使用converters参数指定列转换逻辑
通过converters参数单独指定第二列(索引为1)的转换规则,将其保留为字符串:
import numpy as np values = np.genfromtxt( universal_path, skip_header=1, encoding='unicode_escape', delimiter=None, # 对第2列(索引1)进行字符串转换 converters={1: lambda x: x.decode('unicode_escape')} ) # 查看第二列数据 print(values[:, 1])
问题原因说明
- 原代码中
delimiter=""可替换为delimiter=None,让numpy自动识别任意空白(空格、制表符等)作为分隔符,更适配文件的分隔格式。 - 直接使用
[int, str, ...]作为dtype会被numpy解析为一维数组的元素类型,而非每列的类型,必须使用结构化dtype来定义多列的混合类型。
内容的提问来源于stack exchange,提问作者Camille
相关产品推荐
相关产品推荐

