如何拆分DataFrame多GO条目至独立行并调换列顺序?
Solution
You can achieve this using pandas' string splitting and row explosion capabilities. Here's the complete code:
import pandas as pd # Sample DataFrame (replace with your actual df) data = { 'gene_id': ['LOC_Os02g43120.1', 'LOC_Os12g21850.1', 'LOC_Os06g30330.1', 'LOC_Os07g37690.1'], 'GO': ['GO:0008270 GO:0005515', 'GO:0003700 GO:0006355 GO:0005515 GO:0034645 GO:0043565', 'GO:0005488', 'GO:0016758 GO:0008152'] } df = pd.DataFrame(data) # Step 1: Split the GO column into a list of individual GO terms df['GO'] = df['GO'].str.split() # Step 2: Explode the list to create a row for each GO term df_exploded = df.explode('GO') # Step 3: Reorder columns to place GO first, then gene_id result = df_exploded[['GO', 'gene_id']] # Print the result in the desired format (without index) print(result.to_string(index=False))
Explanation:
- Split the GO column:
str.split()breaks each space-separated GO string into a list of individual terms. - Explode rows:
explode('GO')converts each element in the list into its own row, preserving the corresponding gene_id. - Reorder columns: By selecting
['GO', 'gene_id'], we swap the column order to match your desired output. - Print without index:
to_string(index=False)removes the default pandas index when printing, matching the clean format you provided.
Running this code will produce exactly the output you requested.
内容的提问来源于stack exchange,提问作者sumitra sivaprakasam
相关产品推荐
相关产品推荐

