You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas对列应用返回字典的函数生成多列的解决方案

解决方案

你只需要将apply返回的字典序列转为DataFrame,再和原表按索引拼接即可,以下是两种可行写法:

简易写法,适合小数据量:

# 直接对目标列应用函数后,拆分字典为多列,再拼回原DataFrame
df_result = df_dna.join(df_dna['Forward_Sequence'].apply(freqcount).apply(pd.Series))

针对48万行大样本优化的高效写法,避免多次apply带来的性能损耗:

# 1. 先得到所有行的频率计算结果
freq_result = df_dna['Forward_Sequence'].apply(freqcount)
# 2. 将字典列表直接转为DataFrame,自动匹配字典键为列名
freq_df = pd.DataFrame(freq_result.to_list(), index=freq_result.index)
# 3. 按索引拼回原表
df_result = df_dna.join(freq_df)

原理解释

  • df_dna['Forward_Sequence'].apply(freqcount) 会得到一个元素为字典的Series,索引和原DataFrame完全对齐
  • 用字典列表构造DataFrame时,pandas会自动将每个字典的键识别为列名,字典的值对应填充到对应行的对应列
  • 最后用join方法按索引拼接,不会打乱原数据的行顺序

可选性能优化

如果计算速度达不到要求,可以优化freqcount函数的统计逻辑,用collections.Counter批量统计替代多次count调用,减少字符串遍历次数:

from collections import Counter

def freqcount_optimized(s):
    bases = "".join(s.split("[CG]"))
    total = len(bases)
    single_count = Counter(bases)
    dinuc_count = Counter(bases[i:i+2] for i in range(len(bases)-1))
    outdic = {}
    for b1 in ["A", "G", "C", "T"]:
        outdic[b1] = single_count.get(b1, 0)/total
        for b2 in ["A", "G", "C", "T"]:
            outdic[b1+b2] = dinuc_count.get(b1+b2, 0)/total
    return outdic

内容的提问来源于stack exchange,提问作者charelf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 05:39:05