You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用np.genfromtxt读取含hh:mm:ss格式时间列的文本文件?

问题描述

有如下格式的文本文件:

index   timestamp   polarisation current (A)    signal (V)  head temperature (°C)   head relat.humidity (%RH)   MUGS temperature (°C)   laser voltage (V)   laser current (A)   driver temperature (°C)
0   16:11:24    0.4 0.0006019   26.51   43.5    32.0    11.37   0.3922  26.5
1   16:11:29    0.402   0.0006286   26.51   43.5    32.5    11.41   0.3972  31.5
2   16:11:34    0.404   0.0005828   26.51   43.5    32.5    11.42   0.4048  32.5
3   16:11:38    0.406   0.0006139   26.51   43.5    32.5    11.39   0.3984  32.5

使用以下代码读取文件时,第二列(timestamp)全为NaN:

with open(universal_path,'rt'): 
        values = np.genfromtxt( universal_path, delimiter="", skip_header = 1, encoding='unicode_escape')

尝试设置dtype=None仅能读取第一列;设置dtype = [int, str, float, float, float, float, float, float, float, float]同样仅能读取第一列,需要实现将第二列以字符串形式正确读取。

解决方案

方法1:使用结构化数据类型(推荐)

numpy需要通过结构化数据类型明确指定每一列的类型,而非简单的类型列表。示例代码:

import numpy as np

# 定义对应每一列的结构化数据类型
data_type = np.dtype([
    ('index', int),
    ('timestamp', 'U8'),  # U8表示长度为8的Unicode字符串,适配"HH:MM:SS"格式
    ('polarisation_current', float),
    ('signal', float),
    ('head_temperature', float),
    ('head_humidity', float),
    ('mugs_temperature', float),
    ('laser_voltage', float),
    ('laser_current', float),
    ('driver_temperature', float)
])

# 读取文件,delimiter=None让numpy自动识别任意空白分隔符
values = np.genfromtxt(
    universal_path,
    dtype=data_type,
    skip_header=1,
    encoding='unicode_escape',
    delimiter=None
)

# 访问第二列的时间数据
print(values['timestamp'])

方法2:使用converters参数指定列转换逻辑

通过converters参数单独指定第二列(索引为1)的转换规则,将其保留为字符串:

import numpy as np

values = np.genfromtxt(
    universal_path,
    skip_header=1,
    encoding='unicode_escape',
    delimiter=None,
    # 对第2列(索引1)进行字符串转换
    converters={1: lambda x: x.decode('unicode_escape')}
)

# 查看第二列数据
print(values[:, 1])

问题原因说明

  • 原代码中delimiter=""可替换为delimiter=None,让numpy自动识别任意空白(空格、制表符等)作为分隔符,更适配文件的分隔格式。
  • 直接使用[int, str, ...]作为dtype会被numpy解析为一维数组的元素类型,而非每列的类型,必须使用结构化dtype来定义多列的混合类型。

内容的提问来源于stack exchange,提问作者Camille

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 00:12:05