R语言:基于其他行列值条件删除DataFrame行
Got it, let's tackle this problem step by step. You need to eliminate rows where the REF column is '-', and there's a matching preceding row (same CHR, END, START values) where that row's REF equals the current row's ALT, and its ALT is '-'. Here are two reliable approaches:
1. General Approach (Works regardless of row order)
This method works even if complementary rows aren't adjacent, by grouping rows with identical CHR, END, START values and checking for matching pairs.
First, let's set up your sample data:
import pandas as pd # Sample DataFrame data = { 'CHR': [1, 1, 3, 3], 'END': [1445, 1445, 2787, 2787], 'START': [1446, 1446, 2787, 2787], 'REF': ['G', 'A', 'T', '-'], 'ALT': ['A', 'G', '-', 'T'] } df = pd.DataFrame(data)
Now define a group-wise filtering function to identify and remove the unwanted rows:
# Group by the shared columns grouped = df.groupby(['CHR', 'END', 'START']) def filter_complementary_rows(group): # Isolate rows where REF is '-' dash_ref_rows = group[group['REF'] == '-'] if dash_ref_rows.empty: return group # Collect indices of rows to delete indices_to_drop = [] for idx, row in dash_ref_rows.iterrows(): # Check if there's a matching row with REF = current ALT and ALT = '-' matching_row = group[(group['REF'] == row['ALT']) & (group['ALT'] == '-')] if not matching_row.empty: indices_to_drop.append(idx) # Return the group without the unwanted rows return group.drop(indices_to_drop) # Apply the filter and reset the index cleaned_df = grouped.apply(filter_complementary_rows).reset_index(drop=True)
2. Optimized Approach (Leverages your row order note)
Since you mentioned the rows to delete always come after their matching pair, we can use shift() to compare each row directly with the previous one in its group, which is faster:
# Add a flag to mark rows for deletion df['to_delete'] = grouped.apply( lambda x: (x['REF'] == '-') & (x['ALT'] == x['REF'].shift()) & (x['ALT'].shift() == '-') ).reset_index(drop=True) # Filter out marked rows and clean up the helper column cleaned_df = df[~df['to_delete']].drop('to_delete', axis=1)
Result
Both methods will give you the desired output:
CHR END START REF ALT 0 1 1445 1446 G A 1 1 1445 1446 A G 2 3 2787 2787 T -
内容的提问来源于stack exchange,提问作者Javier bertol chorro

