Python新手求助:用for循环与if判断过滤考勤记录中的重复日期
Hey there! As someone new to Python, this kind of data cleaning is super common, so let's walk through how to solve it using a for loop and if checks exactly like you asked.
First, let's recap your problem: you have a list of datetime objects where some dates are duplicated (like April 5th has two entries), and you want to keep only one record per date, ignoring the rest.
Step-by-Step Solution
The key idea here is to track which dates we've already processed as we loop through the list. We'll use a set for this because checking if an item exists in a set is fast and efficient. Then, for each attendance entry, we'll check if its date has been seen before—if not, we'll keep the entry and mark the date as seen.
Here's the code:
import datetime # Your original attendance data attendance = [ datetime.datetime(2018, 4, 1, 7, 01), datetime.datetime(2018, 4, 2, 7, 01), datetime.datetime(2018, 4, 3, 7, 01), datetime.datetime(2018, 4, 4, 7, 01), datetime.datetime(2018, 4, 5, 7, 01), datetime.datetime(2018, 4, 5, 7, 21), datetime.datetime(2018, 4, 6, 7, 01), datetime.datetime(2018, 4, 7, 7, 01), ] # Initialize a set to track dates we've already added seen_dates = set() # Initialize an empty list to hold our deduplicated results unique_attendance = [] for entry in attendance: # Extract just the date part (ignoring time) entry_date = entry.date() # Check if we haven't seen this date yet if entry_date not in seen_dates: # Add the entry to our results unique_attendance.append(entry) # Mark the date as seen so we skip duplicates later seen_dates.add(entry_date) # Print the result for item in unique_attendance: print(item)
What This Does
entry.date()converts the full datetime object into adateobject, which only includes year/month/day—this is how we compare if two entries are from the same date, regardless of time.- The
seen_datesset keeps track of every date we've already added tounique_attendance. If an entry's date is already in the set, we skip it. - We're keeping the first occurrence of each date in the original list. If you wanted to keep the last occurrence instead, you could adjust the logic (but your question says "ignore other records, only show one per date," so first occurrence works perfectly).
Sample Output
When you run this code, you'll get:
2018-04-01 07:01:00 2018-04-02 07:01:00 2018-04-03 07:01:00 2018-04-04 07:01:00 2018-04-05 07:01:00 2018-04-06 07:01:00 2018-04-07 07:01:00
Notice the duplicate April 5th entry (7:21) is gone—we only keep the first one from that date.
If you have any follow-up questions (like how to adjust for keeping the latest entry instead), feel free to ask!
内容的提问来源于stack exchange,提问作者Dicky Raambo

