如何在Python的Pandas中提取重复行名对应的第二行数据
解决重复表头Excel的数据读取问题
你的问题根源是目标Excel包含两行重复的表头,第一行是无效的,第二行才是真正的列名。直接用默认参数读取会把第一行当作表头,导致后续无法正确获取你需要的第二行对应的数据。
最简单的解决方案:指定表头行
在读取Excel时,直接告诉pandas使用第二行(索引为1,从0开始计数)作为表头,自动跳过第一行无效内容:
df = pd.read_excel(aisc_excel_file, header=1) with open(profiles_lis_file, 'w') as profiles_lis: profiles_lis.write("PROFILE\tWIDTH\tHEIGHT\n") for index, row in df.iterrows(): profile_name = row['AISC_Manual_Label'] width = row['h, in'] height = row['b, in'] profiles_lis.write(f"{profile_name}\t{width}\t{height}\n")
备选方案:手动重置表头(针对特殊格式Excel)
如果上述方法因Excel格式异常不生效,可以手动重置表头:
# 先读取所有内容,不指定表头 df_raw = pd.read_excel(aisc_excel_file, header=None) # 提取第三行及以后的数据(跳过前两行),并将第二行设为列名 df = df_raw.loc[2:] df.columns = df_raw.iloc[1] # 重置索引避免混乱 df = df.reset_index(drop=True) # 后续写入文件代码不变 with open(profiles_lis_file, 'w') as profiles_lis: profiles_lis.write("PROFILE\tWIDTH\tHEIGHT\n") for index, row in df.iterrows(): profile_name = row['AISC_Manual_Label'] width = row['h, in'] height = row['b, in'] profiles_lis.write(f"{profile_name}\t{width}\t{height}\n")
说明
你之前尝试的df.duplicated()是用来检测数据行重复的方法,而你的问题属于表头行重复的列名定义问题,因此该方法无法解决你的需求。
内容的提问来源于stack exchange,提问作者Jude Austin
相关产品推荐
相关产品推荐

