如何高效将HH:MM:SS转换为十进制小时用于gnuplot绘图?
高效将HH:MM:SS转换为十进制小时的方法
你的bash脚本处理慢的核心原因是:循环中反复启动多个外部进程(sed、bc等),且sed -i每次都会重写整个文件,IO和进程启动的开销累加后导致9000条数据耗时42秒。以下是几种高效的替代方案:
1. Awk 实现(性能最优)
Awk是处理文本的原生工具,单进程完成所有处理,无需频繁启动外部程序,效率最高:
awk '{ split($4, time_parts, ":"); decimal_hour = time_parts[1] + time_parts[2]/60 + time_parts[3]/3600; printf "%s %s %s %.4f\n", $1, $2, $3, decimal_hour; }' /tmp/fn > /tmp/fn_converted
执行后,转换结果会写入/tmp/fn_converted,若需替换原文件可执行mv /tmp/fn_converted /tmp/fn。
2. Python 实现
适合熟悉Python的场景,代码可读性强,性能同样出色:
with open('/tmp/fn', 'r') as in_file, open('/tmp/fn_converted', 'w') as out_file: for line in in_file: stripped_line = line.strip() if not stripped_line: continue parts = stripped_line.split() hh, mm, ss = parts[3].split(':') dec_hour = int(hh) + int(mm)/60 + int(ss)/3600 out_file.write(f"{parts[0]} {parts[1]} {parts[2]} {dec_hour:.4f}\n")
3. Perl 实现
Perl处理文本的效率接近Awk,代码简洁紧凑:
open(my $input_fh, '<', '/tmp/fn') or die "无法打开输入文件: $!"; open(my $output_fh, '>', '/tmp/fn_converted') or die "无法打开输出文件: $!"; while (my $line = <$input_fh>) { chomp $line; next if $line =~ /^\s*$/; my ($year, $mon, $day, $time_str) = split /\s+/, $line; my ($h, $m, $s) = split /:/, $time_str; my $dec_hour = $h + $m/60 + $s/3600; printf $output_fh "%s %s %s %.4f\n", $year, $mon, $day, $dec_hour; } close $input_fh; close $output_fh;
效果对比
以上三种方案处理9000条数据的耗时通常都在0.1秒以内,远低于原bash脚本的42秒,完全能应对更大规模的数据集。
内容的提问来源于stack exchange,提问作者dwarnok
相关产品推荐
相关产品推荐

