使用Pandas按日期筛选时为何出现哈希表错误及KeyError: 'Tested On'?
Hey there! Let's work through your Pandas issue step by step. You're hitting two problems: a hash table error when filtering by date, and a KeyError: 'Tested On'. Let's break down why this might be happening and how to fix it, using your sample data as a reference.
First: Diagnose the KeyError: 'Tested On'
This error means Pandas can't find a column named exactly Tested On in your DataFrame. Here are the most likely causes and fixes:
1. Column Name Typos or Hidden Whitespace
Sometimes when importing data (especially from Excel/CSV with merged cells), column names can pick up extra spaces or have inconsistent capitalization. First, check exactly what columns are in your DataFrame:
import pandas as pd # After importing your data print("Current columns:", df.columns.tolist()) print("\nColumn data types:", df.dtypes)
If you see something like ' Tested On ' (with leading/trailing spaces) or 'tested on', clean up the column names:
# Remove whitespace from all column names df.columns = df.columns.str.strip() # Or force consistent capitalization if needed df.columns = df.columns.str.title()
2. Import Issues with Merged Cells
Your sample data has merged cells in the Plan column (e.g., P0 spans two rows). When importing this from Excel, Pandas might read the Plan column with NaN values for the second row, but this shouldn't affect the Tested On column name—unless the import process messed up the header row. Double-check that you're using the right import parameters:
# If importing from Excel df = pd.read_excel("your_file.xlsx", header=0) # Ensure header is set to row 0 # If importing from CSV df = pd.read_csv("your_file.csv", header=0)
Next: Fix the Hash Table Error When Filtering by Date
Hash table errors during date filtering almost always stem from the Tested On column being stored as a string (object type) instead of a datetime type. Pandas can't efficiently hash string dates for filtering, leading to errors. Here's how to fix it:
1. Convert the Date Column to Datetime Type
First, convert Tested On to a proper datetime column. Use pd.to_datetime with errors='coerce' to handle any invalid dates (they'll become NaT, "Not a Time"):
df['Tested On'] = pd.to_datetime(df['Tested On'], errors='coerce')
Verify the conversion worked by checking df.dtypes—Tested On should now show datetime64[ns].
2. Filter Dates Correctly
Now you can filter without hash table issues. For example, to get all rows tested on 2018-08-11:
target_date = pd.to_datetime('2018-08-11') filtered_data = df[df['Tested On'] == target_date]
Or to filter a date range (e.g., between 2018-08-01 and 2018-08-10):
start = pd.to_datetime('2018-08-01') end = pd.to_datetime('2018-08-10') filtered_data = df[(df['Tested On'] >= start) & (df['Tested On'] <= end)]
3. Upgrade Pandas (If Needed)
Occasionally, hash table bugs are fixed in newer Pandas versions. If you're running an older version, try upgrading:
pip install --upgrade pandas
Testing with Your Sample Data
Using your provided data, after fixing the column name and converting the date column, filtering should work as expected. For example, filtering for 2018-08-11 would return the first and fourth rows from your sample.
内容的提问来源于stack exchange,提问作者gomathi subramanian

