DataFrame行复制异常:多字符category列未按预期拆分
问题解决:拆分DataFrame的category列并复制对应行
需求说明
处理指定结构的DataFrame:
- 当
category列值为多字符(如MV)时,拆分每个字符并复制该行,每行对应一个拆分后的字符 - 当
category列值为NO_COG_HIT时,直接保留该行
测试数据
og cogs consensus p_function category gene_1 gene_2 t dnds dn ds OG0000190 COG0593 99 / 99 Chromosomal replication initiation ATPase DnaA (DnaA) L apS2gMQ_00001 GAWBhUD_01925 0.0194 0.126 0.0021 0.0163 OG0000190 COG0593 99 / 99 Chromosomal replication initiation ATPase DnaA (DnaA) L apS2gMQ_00001 GcPBA0T_00001 0.0174 0.001 0 0.0168 OG0000335 COG0845 99 / 99 Multidrug efflux pump subunit AcrA (membrane-fusion protein) (AcrA) MV JdgVjSO_00092 IlhnQ8K_01601 0.0244 0.1508 0.0027 0.0181 OG0000335 COG0845 99 / 99 Multidrug efflux pump subunit AcrA (membrane-fusion protein) (AcrA) MV JdgVjSO_00092 IsqAoZB_00822 0.0359 0.1083 0.0029 0.0265 OG0000532 COG0534 99 / 99 Na+-driven multidrug efflux pump, DinF/NorM/MATE family (NorM) V pr2jcFN_01326 528cT6K_01654 0.1306 0.1176 0.013 0.1105 OG0000532 COG0534 99 / 99 Na+-driven multidrug efflux pump, DinF/NorM/MATE family (NorM) V GcPBA0T_00567 7QtjQYC_01559 0.0502 0.1786 0.0067 0.0373 OG0000223 2DSC2 99 / 99 NO_COG_HIT NO_COG_HIT HQyC1X2_00055 BDcxYt7_01158 0.0083 99 0.0053 1e-04 OG0000223 2DSC2 99 / 99 NO_COG_HIT NO_COG_HIT kNAVz3k_01037 7QtjQYC_00282 0.0083 99 0.0053 1e-04
当前问题
现有脚本能复制行,但未将category拆分为单个字符,输出中category仍为原多字符值。
错误原因
- 循环拆分字符时,
category字段赋值为原行的test_file.at[r, 'category'],而非拆分后的单个字符letter else分支中cog = letter会引发未定义错误,因为letter仅在if分支的循环内存在
修正后的代码
基础修复版本(保留原逻辑)
import pandas as pd # 加载数据 test_file = pd.read_csv("/path/test.tsv", sep="\t", names=['og', 'cogs', 'consensus', 'p_function', 'category', 'gene_1', 'gene_2', 't', 'dnds', 'dn', 'ds']) new_test_file = pd.DataFrame(columns=test_file.columns) for r in test_file.index: join_cog = test_file.at[r, 'category'] if join_cog != 'NO_COG_HIT': # 遍历每个拆分后的字符 for letter in join_cog: # 复制原行数据,替换category为当前字符 df_tmp = test_file.loc[[r]].copy() df_tmp['category'] = letter new_test_file = pd.concat([new_test_file, df_tmp], ignore_index=True) else: # 直接保留原行 new_test_file = pd.concat([new_test_file, test_file.loc[[r]]], ignore_index=True) # 输出结果 print(new_test_file)
高效优化版本(利用pandas内置方法,避免循环)
import pandas as pd # 加载数据 test_file = pd.read_csv("/path/test.tsv", sep="\t", names=['og', 'cogs', 'consensus', 'p_function', 'category', 'gene_1', 'gene_2', 't', 'dnds', 'dn', 'ds']) # 处理category列:将非NO_COG_HIT的值拆分为字符列表,NO_COG_HIT保持原样 test_file['category'] = test_file['category'].apply( lambda x: list(x) if x != 'NO_COG_HIT' else [x] ) # 展开列表,生成多行 new_test_file = test_file.explode('category', ignore_index=True) # 输出结果 print(new_test_file)
验证结果
修正后,category为MV的行将拆分为两行,分别对应M和V,符合预期输出。
内容的提问来源于stack exchange,提问作者Someone_1313
相关产品推荐
相关产品推荐

