Pandas对列应用返回字典的函数生成多列的解决方案
解决方案
你只需要将apply返回的字典序列转为DataFrame,再和原表按索引拼接即可,以下是两种可行写法:
简易写法,适合小数据量:
# 直接对目标列应用函数后,拆分字典为多列,再拼回原DataFrame df_result = df_dna.join(df_dna['Forward_Sequence'].apply(freqcount).apply(pd.Series))
针对48万行大样本优化的高效写法,避免多次apply带来的性能损耗:
# 1. 先得到所有行的频率计算结果 freq_result = df_dna['Forward_Sequence'].apply(freqcount) # 2. 将字典列表直接转为DataFrame,自动匹配字典键为列名 freq_df = pd.DataFrame(freq_result.to_list(), index=freq_result.index) # 3. 按索引拼回原表 df_result = df_dna.join(freq_df)
原理解释
df_dna['Forward_Sequence'].apply(freqcount)会得到一个元素为字典的Series,索引和原DataFrame完全对齐- 用字典列表构造DataFrame时,pandas会自动将每个字典的键识别为列名,字典的值对应填充到对应行的对应列
- 最后用
join方法按索引拼接,不会打乱原数据的行顺序
可选性能优化
如果计算速度达不到要求,可以优化freqcount函数的统计逻辑,用collections.Counter批量统计替代多次count调用,减少字符串遍历次数:
from collections import Counter def freqcount_optimized(s): bases = "".join(s.split("[CG]")) total = len(bases) single_count = Counter(bases) dinuc_count = Counter(bases[i:i+2] for i in range(len(bases)-1)) outdic = {} for b1 in ["A", "G", "C", "T"]: outdic[b1] = single_count.get(b1, 0)/total for b2 in ["A", "G", "C", "T"]: outdic[b1+b2] = dinuc_count.get(b1+b2, 0)/total return outdic
内容的提问来源于stack exchange,提问作者charelf
相关产品推荐
相关产品推荐

