Pandas:删除Time为0的行及同ID的该行之前所有行
Solution to Remove Target Rows Based on Time=0 and ID
Got it, let's break down how to solve this problem effectively. Your goal is to:
- Delete all rows that share the same ID as any row where
Timeequals 0, but only those rows that appear before the Time=0 row - Delete the Time=0 row itself
Step-by-Step Implementation with Pandas
Since this is tabular data, using Python's Pandas library is the most straightforward approach. Here's a complete, tested solution:
import pandas as pd # Replace this with your actual data loading (e.g., pd.read_csv("your_data.csv")) data = { 'Time': ['2018-04-23 08:09:52.558', '2018-04-23 14:26:12.294', '2018-04-23 15:19:47.667', '2018-04-23 18:05:25.417', '2018-04-23 18:58:20.776', '2018-04-23 18:59:50.883', '2018-04-23 19:29:30.500', '0'], # Marked the target Time=0 row 'Author': [60, 360, 45, 240, 180, 120, 300, 0], 'ID': [1744025, 1244021, 1244021, 1249997, 2185555, 2185555, 1686620, 1249997], 'Date': [44, 10, 10, 19, 19, 19, 19, 19] } df = pd.DataFrame(data) # 1. Identify all rows where Time equals 0 zero_time_rows = df[df['Time'] == '0'] # 2. Collect indices of rows to delete rows_to_drop = [] for idx, target_row in zero_time_rows.iterrows(): target_id = target_row['ID'] # Add all rows with the same ID that come BEFORE the Time=0 row prior_matching_rows = df[(df['ID'] == target_id) & (df.index < idx)] rows_to_drop.extend(prior_matching_rows.index.tolist()) # Add the Time=0 row itself rows_to_drop.append(idx) # 3. Remove duplicate indices (in case of overlapping matches) and clean the DataFrame df_cleaned = df.drop(list(set(rows_to_drop))).reset_index(drop=True) # Print the result print(df_cleaned)
What This Code Does
- Data Loading: We start by defining sample data (replace this with your actual data source like a CSV file).
- Target Row Identification: We first find all rows where
Timeis 0. - Collect Rows to Delete: For each Time=0 row, we find all rows with the same ID that appear earlier in the dataset, plus the Time=0 row itself.
- Clean the Data: We drop all collected indices (removing duplicates to avoid errors) and reset the index for readability.
Example Output
Running this code on the sample data will produce:
Time Author ID Date 0 2018-04-23 08:09:52.558 60 1744025 44 1 2018-04-23 14:26:12.294 360 1244021 10 2 2018-04-23 15:19:47.667 45 1244021 10 3 2018-04-23 18:58:20.776 180 2185555 19 4 2018-04-23 18:59:50.883 120 2185555 19 5 2018-04-23 19:29:30.500 300 1686620 19
Notice that the row with ID 1249997 (both the prior row and the Time=0 row) have been removed, as required.
Notes
- If your
Timecolumn is numeric instead of string, adjust the condition todf['Time'] == 0instead of'0'. - This solution handles multiple Time=0 rows (each with different IDs) correctly—each will trigger deletion of their prior matching ID rows and themselves.
内容的提问来源于stack exchange,提问作者tbienias
相关产品推荐
相关产品推荐

