如何用循环将CSV中多行IVC数据转换为多列Voltage/Current格式?
处理多组IV数据的CSV整理方案
针对你描述的CSV文件格式,我会用Python实现自动遍历并整理多组Voltage/Current数据到独立列的需求。下面分两种方案(纯CSV模块和Pandas,后者更简洁高效)来讲解:
方案一:使用Pandas(推荐,适合大型数据集)
Pandas处理这类结构化数据转换非常方便,代码逻辑清晰,也能轻松处理你提到的大量数据(到第7117行及后续组)。
步骤说明:
- 跳过前91行,读取从第92行开始的所有内容
- 识别每个数据组的标题(比如
Voltage:、Current:、Voltage 2:等) - 收集每个标题对应的所有数值,整理成字典(键为列名,值为数值列表)
- 将字典转换为DataFrame,自动对齐行,缺失值补NaN
- 保存为新的CSV文件
代码实现:
import pandas as pd # 读取文件,跳过前91行(索引从0开始,第92行对应索引91) with open('your_file.csv', 'r') as f: lines = [line.strip() for line in f.readlines()[91:]] data_dict = {} current_key = None current_values = [] for line in lines: if not line: # 跳过空行 continue # 判断是否是组标题(包含"Voltage"或"Current") if 'Voltage' in line or 'Current' in line: # 如果之前有正在收集的组,先保存到字典 if current_key is not None: # 处理列名:比如"Voltage 2:" → "Voltage1","Current:" → "Current" clean_key = current_key.replace(':', '').strip() if '2' in clean_key: clean_key = clean_key.replace('2', '1') elif '3' in clean_key: clean_key = clean_key.replace('3', '2') # 若有更多组(如Voltage 4),可继续添加elif分支调整编号 data_dict[clean_key] = current_values # 开始新的组 current_key = line current_values = [] else: # 收集数值,转成float类型 try: current_values.append(float(line)) except ValueError: # 遇到非数值行(如无关内容)则跳过 continue # 处理最后一个未保存的组 if current_key is not None: clean_key = current_key.replace(':', '').strip() if '2' in clean_key: clean_key = clean_key.replace('2', '1') elif '3' in clean_key: clean_key = clean_key.replace('3', '2') data_dict[clean_key] = current_values # 转换为DataFrame,自动对齐不同长度的组数据 df = pd.DataFrame(dict([(k, pd.Series(v)) for k, v in data_dict.items()])) # 保存整理后的CSV df.to_csv('cleaned_iv_data.csv', index=False)
关键细节:
- 列名转换逻辑可根据实际组标识调整(比如如果是
Voltage A而非Voltage 2,可修改clean_key的生成规则) - 自动跳过空行和非数值行,避免处理错误数据
- 不同长度的组数据会自动对齐,缺失位置填充
NaN,保证每行对应一个数据点的所有组测量值
方案二:使用Python内置CSV模块(无需额外安装库)
如果你不想安装Pandas,用Python自带的csv模块也能实现,适合环境受限的场景:
代码实现:
import csv input_file = 'your_file.csv' output_file = 'cleaned_iv_data.csv' # 读取并预处理数据 with open(input_file, 'r') as f: lines = [line.strip() for line in f.readlines()[91:]] data_groups = {} current_group = None current_values = [] for line in lines: if not line: continue if 'Voltage' in line or 'Current' in line: if current_group is not None: # 清理组名并保存 clean_name = current_group.replace(':', '').strip() if '2' in clean_name: clean_name = clean_name.replace('2', '1') data_groups[clean_name] = current_values.copy() current_group = line current_values = [] else: try: current_values.append(float(line)) except ValueError: continue # 处理最后一个组 if current_group is not None: clean_name = current_group.replace(':', '').strip() if '2' in clean_name: clean_name = clean_name.replace('2', '1') data_groups[clean_name] = current_values # 找到最长的数值列表,确定总行数 max_rows = max(len(v) for v in data_groups.values()) # 填充每个组的列表到统一长度,不足补空字符串 for key in data_groups: while len(data_groups[key]) < max_rows: data_groups[key].append('') # 准备写入CSV的内容:表头 + 每行数据 headers = list(data_groups.keys()) rows = [] for i in range(max_rows): row = [data_groups[key][i] for key in headers] rows.append(row) # 写入文件 with open(output_file, 'w', newline='') as f: writer = csv.writer(f) writer.writerow(headers) writer.writerows(rows)
使用提示:
- 请把代码中的
your_file.csv替换成你实际的文件名 - 如果组的命名规则和示例不同(比如用字母编号),请调整列名转换的逻辑
- 若数据中有特殊格式的异常行,可根据实际情况修改跳过规则
内容的提问来源于stack exchange,提问作者hebin manuel
相关产品推荐
相关产品推荐

