如何修改Python代码以仅对DataFrame中的有效日期计算距今天数
Hey, let's work through this problem together. Your code is hitting an error because when it encounters the "N/A" value in the Date column, datetime.strptime can't parse it into a valid date. On top of that, your loop is actually iterating over the DataFrame's column names (not rows), which would cause additional issues even without the invalid date. Here are two solid solutions:
Solution 1: Fix Your Loop with Error Handling
If you want to stick with a loop-based approach (great for learning!), we'll correct the row iteration and add checks for invalid dates:
import datetime import pandas as pd newdft = [] # Use iterrows() to loop through each row (idx = row index, row = row data) for idx, row in dft.iterrows(): temp_row = row.copy() date_value = row["Date"] # Check if the date is not "N/A" and not empty/missing if pd.notna(date_value) and date_value != "N/A": try: # Parse the date string parsed_date = datetime.datetime.strptime(date_value, "%m/%d/%Y") # Calculate days between current date and parsed date days_diff = (datetime.datetime.now() - parsed_date).days temp_row["Days"] = days_diff except ValueError: # Catch any other invalid date formats (e.g., typos like "13/01/2022") temp_row["Days"] = pd.NA else: # Handle "N/A" or missing values temp_row["Days"] = pd.NA newdft.append(temp_row) # Convert the list of rows back to a DataFrame newdft = pd.DataFrame(newdft)
Solution 2: Use Pandas Native Methods (Recommended for Efficiency)
Pandas has built-in tools for date handling that are way faster than manual loops, especially with large datasets. Here's the cleaner approach:
- First, convert the
Datecolumn to proper datetime type, forcing invalid values toNaT(Pandas' "Not a Time" marker):
dft["Date"] = pd.to_datetime(dft["Date"], format="%m/%d/%Y", errors="coerce")
The errors="coerce" parameter automatically turns unparseable values (like "N/A") into NaT.
- Calculate the days difference in one line—
NaTvalues will result inNaNfor theDayscolumn:
dft["Days"] = (datetime.datetime.now() - dft["Date"]).dt.days
Bonus: Customize Missing Values
If you want to replace NaN in the Days column with a default value (like 0), just add:
dft["Days"] = dft["Days"].fillna(0)
内容的提问来源于stack exchange,提问作者Mary

