如何用Python更高效地重组CSV文件?(无Pandas环境)
无Pandas实现CSV重组方案
直接用Python标准库csv模块就能高效完成需求,代码精简且无需额外依赖:
import csv input_file = "etmgeg_380_test.csv" output_file = "formatted_rh_data.csv" # 定位表头并处理数据 with open(input_file, 'r', newline='', encoding='utf-8') as infile: # 跳过# STN开头之前的所有行 for line in infile: stripped_line = line.strip() if stripped_line.startswith('# STN'): # 处理表头:移除开头的#和空格,分割列名 header = [col.strip() for col in stripped_line.lstrip('#').split(',')] break # 获取目标列的索引 date_col_idx = header.index('YYYYMMDD') rh_col_idx = header.index('RH') # 写入格式化后的数据 with open(output_file, 'w', newline='', encoding='utf-8') as outfile: writer = csv.writer(outfile) # 写入适配InfluxDB的表头 writer.writerow(['time', 'rh']) # 逐行处理数据 reader = csv.reader(infile) for row in reader: if not row: continue # 格式化日期:YYYYMMDD → YYYY-MM-DD raw_date = row[date_col_idx] formatted_date = f"{raw_date[:4]}-{raw_date[4:6]}-{raw_date[6:]}" # 提取RH值 rh_value = row[rh_col_idx] writer.writerow([formatted_date, rh_value])
关键细节说明:
- 流式读写:逐行处理数据,内存占用极低,适合大文件场景
- 快速日期格式化:用字符串切片直接转换格式,比
datetime解析更高效(若日期格式固定) - 动态列定位:通过表头列表的
index方法获取列位置,避免硬编码索引导致的维护问题
如果需要严格校验日期格式,可替换为datetime模块解析:
from datetime import datetime formatted_date = datetime.strptime(raw_date, '%Y%m%d').strftime('%Y-%m-%d')
内容的提问来源于stack exchange,提问作者Joost
相关产品推荐
相关产品推荐

