You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame行复制异常:多字符category列未按预期拆分

问题解决:拆分DataFrame的category列并复制对应行

需求说明

处理指定结构的DataFrame:

  • 当category列值为多字符(如MV)时,拆分每个字符并复制该行,每行对应一个拆分后的字符
  • 当category列值为NO_COG_HIT时,直接保留该行

测试数据

og  cogs    consensus   p_function  category    gene_1  gene_2  t   dnds    dn  ds
OG0000190   COG0593 99 / 99 Chromosomal replication initiation ATPase DnaA (DnaA)   L   apS2gMQ_00001   GAWBhUD_01925   0.0194  0.126   0.0021  0.0163
OG0000190   COG0593 99 / 99 Chromosomal replication initiation ATPase DnaA (DnaA)   L   apS2gMQ_00001   GcPBA0T_00001   0.0174  0.001   0   0.0168
OG0000335   COG0845 99 / 99 Multidrug efflux pump subunit AcrA (membrane-fusion protein) (AcrA) MV  JdgVjSO_00092   IlhnQ8K_01601   0.0244  0.1508  0.0027  0.0181
OG0000335   COG0845 99 / 99 Multidrug efflux pump subunit AcrA (membrane-fusion protein) (AcrA) MV  JdgVjSO_00092   IsqAoZB_00822   0.0359  0.1083  0.0029  0.0265
OG0000532   COG0534 99 / 99 Na+-driven multidrug efflux pump, DinF/NorM/MATE family (NorM)  V   pr2jcFN_01326   528cT6K_01654   0.1306  0.1176  0.013   0.1105
OG0000532   COG0534 99 / 99 Na+-driven multidrug efflux pump, DinF/NorM/MATE family (NorM)  V   GcPBA0T_00567   7QtjQYC_01559   0.0502  0.1786  0.0067  0.0373
OG0000223   2DSC2   99 / 99 NO_COG_HIT  NO_COG_HIT  HQyC1X2_00055   BDcxYt7_01158   0.0083  99  0.0053  1e-04
OG0000223   2DSC2   99 / 99 NO_COG_HIT  NO_COG_HIT  kNAVz3k_01037   7QtjQYC_00282   0.0083  99  0.0053  1e-04

当前问题

现有脚本能复制行,但未将category拆分为单个字符,输出中category仍为原多字符值。

错误原因

  1. 循环拆分字符时,category字段赋值为原行的test_file.at[r, 'category'],而非拆分后的单个字符letter
  2. else分支中cog = letter会引发未定义错误,因为letter仅在if分支的循环内存在

修正后的代码

基础修复版本(保留原逻辑)

import pandas as pd

# 加载数据
test_file = pd.read_csv("/path/test.tsv",
                        sep="\t", names=['og', 'cogs', 'consensus', 'p_function', 'category', 'gene_1', 'gene_2', 't',
                                         'dnds', 'dn', 'ds'])

new_test_file = pd.DataFrame(columns=test_file.columns)

for r in test_file.index:
    join_cog = test_file.at[r, 'category']
    if join_cog != 'NO_COG_HIT':
        # 遍历每个拆分后的字符
        for letter in join_cog:
            # 复制原行数据,替换category为当前字符
            df_tmp = test_file.loc[[r]].copy()
            df_tmp['category'] = letter
            new_test_file = pd.concat([new_test_file, df_tmp], ignore_index=True)
    else:
        # 直接保留原行
        new_test_file = pd.concat([new_test_file, test_file.loc[[r]]], ignore_index=True)

# 输出结果
print(new_test_file)

高效优化版本(利用pandas内置方法,避免循环)

import pandas as pd

# 加载数据
test_file = pd.read_csv("/path/test.tsv",
                        sep="\t", names=['og', 'cogs', 'consensus', 'p_function', 'category', 'gene_1', 'gene_2', 't',
                                         'dnds', 'dn', 'ds'])

# 处理category列:将非NO_COG_HIT的值拆分为字符列表,NO_COG_HIT保持原样
test_file['category'] = test_file['category'].apply(
    lambda x: list(x) if x != 'NO_COG_HIT' else [x]
)

# 展开列表,生成多行
new_test_file = test_file.explode('category', ignore_index=True)

# 输出结果
print(new_test_file)

验证结果

修正后,category为MV的行将拆分为两行,分别对应M和V,符合预期输出。

内容的提问来源于stack exchange,提问作者Someone_1313

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 03:12:20