You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python更高效地重组CSV文件?(无Pandas环境)

无Pandas实现CSV重组方案

直接用Python标准库csv模块就能高效完成需求,代码精简且无需额外依赖:

import csv

input_file = "etmgeg_380_test.csv"
output_file = "formatted_rh_data.csv"

# 定位表头并处理数据
with open(input_file, 'r', newline='', encoding='utf-8') as infile:
    # 跳过# STN开头之前的所有行
    for line in infile:
        stripped_line = line.strip()
        if stripped_line.startswith('# STN'):
            # 处理表头:移除开头的#和空格,分割列名
            header = [col.strip() for col in stripped_line.lstrip('#').split(',')]
            break
    
    # 获取目标列的索引
    date_col_idx = header.index('YYYYMMDD')
    rh_col_idx = header.index('RH')
    
    # 写入格式化后的数据
    with open(output_file, 'w', newline='', encoding='utf-8') as outfile:
        writer = csv.writer(outfile)
        # 写入适配InfluxDB的表头
        writer.writerow(['time', 'rh'])
        
        # 逐行处理数据
        reader = csv.reader(infile)
        for row in reader:
            if not row:
                continue
            # 格式化日期:YYYYMMDD → YYYY-MM-DD
            raw_date = row[date_col_idx]
            formatted_date = f"{raw_date[:4]}-{raw_date[4:6]}-{raw_date[6:]}"
            # 提取RH值
            rh_value = row[rh_col_idx]
            writer.writerow([formatted_date, rh_value])

关键细节说明:

  • 流式读写:逐行处理数据,内存占用极低,适合大文件场景
  • 快速日期格式化:用字符串切片直接转换格式,比datetime解析更高效(若日期格式固定)
  • 动态列定位:通过表头列表的index方法获取列位置,避免硬编码索引导致的维护问题

如果需要严格校验日期格式,可替换为datetime模块解析:

from datetime import datetime
formatted_date = datetime.strptime(raw_date, '%Y%m%d').strftime('%Y-%m-%d')

内容的提问来源于stack exchange,提问作者Joost

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 02:47:02