Python:如何从多个同结构DataFrame生成指定格式的唯一DataFrame
Python Pandas 整合多DataFrame的四核苷酸频率数据
解决方法
你手里的每个sample_df都是长格式数据,直接转成目标的宽格式后合并所有结果即可,不用管一开始的空DataFrame,这种方式更高效。
1. 单个DataFrame转宽格式
写个简单函数处理单个样本数据,用pivot把长格式转成宽格式:
import pandas as pd # 先定义你提供的示例数据 colnames = ('ACCT', 'CTAT', 'AAAT', 'ATCG')*3 sample_df = pd.DataFrame() sample_df['tetran'] = colnames sample_df['Frequency'] = (423, 512, 25, 123,632,124,614,73,14,75,311,155) conids = ("cl1_42", "cl1_41", "cl2_31") rep_conids = [val for val in conids for _ in range(4)] sample_df['contig_id'] = rep_conids sample_df_2 = pd.DataFrame() sample_df_2['tetran'] = colnames sample_df_2['Frequency'] = (724, 132, 4, 102,423,402,616,734,153,751,31,55) conids_2 = ("se1_51", "se1_21", "se2_53") rep_conids_2 = [val for val in conids_2 for _ in range(4)] sample_df_2['contig_id'] = rep_conids_2 # 处理单个样本的函数 def process_sample(df): # 转成宽格式:contig_id作为索引,tetran作为列,Frequency作为对应值 wide_df = df.pivot(index='contig_id', columns='tetran', values='Frequency') # 按指定顺序排列列(可选,确保列顺序为ACCT/CTAT/AAAT/ATCG) wide_df = wide_df[['ACCT', 'CTAT', 'AAAT', 'ATCG']] return wide_df # 处理两个示例样本 processed1 = process_sample(sample_df) processed2 = process_sample(sample_df_2)
2. 合并所有处理后的DataFrame
用pd.concat把所有处理好的宽格式DataFrame拼接起来:
final_result = pd.concat([processed1, processed2]) # 如果不想让contig_id作为索引,可执行以下代码重置索引: # final_result = final_result.reset_index() print(final_result)
最终输出
运行后得到的结果完全符合你的需求:
| contig_id | ACCT | CTAT | AAAT | ATCG |
|---|---|---|---|---|
| cl1_42 | 423 | 512 | 25 | 123 |
| cl1_41 | 632 | 124 | 614 | 73 |
| cl2_31 | 14 | 75 | 311 | 155 |
| se1_51 | 724 | 132 | 4 | 102 |
| se1_21 | 423 | 402 | 616 | 734 |
| se2_53 | 153 | 751 | 31 | 55 |
补充说明
- 每个
contig_id对应的四个tetran是唯一的,所以用pivot刚好匹配,不需要额外聚合操作。 - 要是有更多的
sample_df,直接把处理后的结果都加到concat的列表里就行,比如pd.concat([processed1, processed2, processed3, ...])。 - 一开始的空DataFrame完全不需要用到,直接处理样本数据再合并比往空表里填值高效得多。
内容的提问来源于stack exchange,提问作者Valentin
相关产品推荐
相关产品推荐

