Pandas导入CSV后用concat合并数据帧出现列偏移问题求助
问题分析与解决方案
问题现象
循环录入员工数据存入DataFrame并导出CSV后,导入CSV进行concat合并时出现列偏移:
- 循环录入两次后导入CSV,数据从第3行开始,还新增列,最终列数变为8列;
- 重启程序先导入CSV正常,但录入第三条数据时,新数据被放到新列中。
核心原因
- CSV导出时默认写入索引列:
df_Employee.to_csv()默认会把DataFrame的索引写入CSV第一列,导入时用header=None会把这列索引当成普通数据列,同时原DataFrame的列名会被当成第一行数据,导致列数多了一列,合并时列对齐失效。 - 导入时未正确读取表头:原DataFrame有明确列名(对应
index1),但导入时用header=None会自动生成数字列名,和新录入数据的列名不匹配,concat时因列名不一致出现偏移、新增列。 - 数据构造与合并的隐性对齐问题:新录入数据通过转置
df.T合并,虽然当前能匹配列名,但一旦列名定义有偏差,就会触发列对齐错误。
解决方法
1. 修正CSV导出逻辑,排除索引列
导出时添加index=False,避免把索引写入CSV:
df_Employee.to_csv("Employee Details.csv", index=False)
2. 正确导入CSV,匹配列名
导入时指定header=0读取原有列名,确保导入的DataFrame列名和新数据一致:
# 预先定义统一列名 index1 = ["Name", "Age", "Gender", "Location", "Designation", "Salary"] # 导入CSV,读取表头作为列名 df = pd.read_csv("Employee Details.csv", header=0) # 合并数据 df_Employee = pd.concat([df_Employee, df], ignore_index=True)
注意:首次启动程序时,要先初始化空DataFrame,避免合并报错:
import pandas as pd index1 = ["Name", "Age", "Gender", "Location", "Designation", "Salary"] df_Employee = pd.DataFrame(columns=index1)
3. 优化新数据录入的构造方式(更直观)
直接构造单行DataFrame,无需转置,减少出错概率:
name = input("Please Enter the Employee's Name:") age = int(input("Please Enter the Employee's Age:")) gender = input("Please Enter the Employee's Gender:") Loc = input("Please Enter the Employee's Location:") Des = input("Please Enter the Employee's Designation:") Sal = int(input("Please Enter the Employee's Salary:")) # 直接生成符合列名的单行数据 new_row = pd.DataFrame([[name, age, gender, Loc, Des, Sal]], columns=index1) df_Employee = pd.concat([df_Employee, new_row], ignore_index=True)
内容的提问来源于stack exchange,提问作者Anish Das
相关产品推荐
相关产品推荐

