从DataFrame提取from/to_location含空值的记录索引及对应列名
Got it, let's work through this problem step by step. First, let's recreate the DataFrame you provided using pandas so we can test our solution directly:
import pandas as pd # Build the sample DataFrame df = pd.DataFrame({ 'flight_id': [1, 2, 4, 9, 3, 4, 9], 'from_location': ['Vancouver', 'Amsterdam', None, 'Halmstad', 'Brisbane', 'Johannesburg', None], 'to_location': ['Toronto', 'Tokyo', 'Glasgow', 'Athens', None, 'Venice', None], 'schedule': ['3-Jan', '15-Feb', '12-Jan', '21-Jan', '4-Feb', '12-Jan', '3-Mar'] })
Next, we need to identify rows where either from_location or to_location is None, then pair each row's index with the corresponding column name(s) that meet the condition. Here's a straightforward, readable way to do this:
# Initialize an empty list to store our result tuples result = [] # Iterate over each row along with its index for idx, row in df.iterrows(): # Check if from_location is missing and add the tuple if true if pd.isna(row['from_location']): result.append((idx, 'from_location')) # Check if to_location is missing and add the tuple if true if pd.isna(row['to_location']): result.append((idx, 'to_location'))
When you run this code, the result variable will hold exactly what you need:
[(2, 'from_location'), (4, 'to_location'), (6, 'from_location'), (6, 'to_location')]
Quick breakdown of the logic:
- We use
pd.isna()instead of checking forNonedirectly because pandas convertsNonevalues in object-type columns toNaN, and this method reliably detects both cases. - We loop through each row with its index, then check each target column individually. For every match, we add a tuple of (row index, column name) to our result list.
- For rows where both columns are
None(like index 6), we add two tuples—one for each column, which aligns with your requirement to pair the index with every qualifying column name.
If you prefer a more concise approach using list comprehensions, you can also write it like this (it produces the exact same result):
result = [ (idx, col) for idx, row in df.iterrows() for col in ['from_location', 'to_location'] if pd.isna(row[col]) ]
Either implementation will give you the list of (index, column name) tuples you're looking for.
内容的提问来源于stack exchange,提问作者Kingz

