Pandas同一列数据筛选问题:保留仅含Pfam分析的蛋白条目并删除同时含Pfam与SMART分析的条目及报错排查
Hey there! Let's break down what's going wrong and how to get the result you want.
First, Why Did You Get That Error?
Your original code tried to use an if statement with a pandas Series (the result of (df_filtered['analysis']=='Pfam')&(df_filtered['analysis']=='SMART')), which doesn't work in pandas. A Series has multiple boolean values, and pandas can't tell if you want to check if all values are True, any value is True, etc.—hence the ValueError about ambiguous truth values.
On top of that, that condition will never be True for any row: a single row's analysis can't be both 'Pfam' and 'SMART' at the same time. So even if the if worked, it wouldn't do what you need.
Your Goal Recap
You want to:
- Remove all rows for proteins that have both 'Pfam' and 'SMART' analyses
- Keep only rows for proteins that have only 'Pfam' analyses
The Solution Code
Assuming your protein ID column is named protein_id (replace this with your actual column name):
# Step 1: Group by protein ID and get all unique analysis types per protein protein_analysis_groups = df_filtered.groupby('protein_id')['analysis'].unique() # Step 2: Identify proteins that have both Pfam and SMART analyses invalid_proteins = protein_analysis_groups[ protein_analysis_groups.apply(lambda x: 'Pfam' in x and 'SMART' in x) ].index # Step 3: Filter out rows for those invalid proteins df_cleaned = df_filtered[~df_filtered['protein_id'].isin(invalid_proteins)]
Alternative: If You Want to Keep Pfam Rows for Proteins That Have Both
If your actual goal was to keep the 'Pfam' rows (but delete 'SMART' rows) for proteins that have both analyses (instead of deleting the entire protein), use this code instead:
# Identify proteins with both Pfam and SMART invalid_proteins = df_filtered.groupby('protein_id')['analysis'].unique().apply( lambda x: 'Pfam' in x and 'SMART' in x ).index # Keep rows that are either from valid proteins (only Pfam) OR are Pfam rows from invalid proteins df_cleaned = df_filtered[ (~df_filtered['protein_id'].isin(invalid_proteins)) | (df_filtered['analysis'] == 'Pfam') ]
How This Works
- We group the DataFrame by protein ID to check which analyses each protein has
- We flag proteins that have both types of analysis
- We filter the original DataFrame to exclude (or selectively keep) rows based on those flags
内容的提问来源于stack exchange,提问作者Aurinko

