如何检查Pandas DataFrame列是否包含指定数值区间的所有值并找出缺失值
Here's an optimized, easy-to-follow approach to solve your problem, with clear handling of edge cases:
Final Code
def prepare_unique_data(self): missing_values = [] for idx, (min_val, max_val) in enumerate(self.min_max): # Skip columns where min is 0 (per your requirement to ignore 0 as minimum) if min_val == 0: missing_values.append([]) continue # Get the current column's data column = self.df_smalltrain.iloc[:, idx] # Extract unique integer values (drop NaNs, convert to int to handle whole-number floats) unique_ints = set(column.dropna().astype(int).unique()) # Generate the complete range of integers we need to check required_range = set(range(min_val, max_val + 1)) # Find values in the required range that are missing from the column missing = sorted(required_range - unique_ints) missing_values.append(missing) # Output or return the result print(missing_values) return missing_values
How It Works
Let’s break down the key steps to understand what’s happening:
Align Columns with Min-Max Pairs: Using
enumerateonself.min_maxensures we match each min-max pair to the correct column in your DataFrame—critical for accurate checks.Ignore 0 as Minimum: If a pair’s min value is 0, we add an empty list to the result (since you specified 0 shouldn’t be treated as a valid minimum).
Extract Valid Unique Integers: For each column, we drop NaN values (they don’t count as integers) and convert all values to integers (to handle cases where your column might have whole-number floats like
2.0). Storing these in a set allows fast lookups.Generate Required Integer Range: We create all integers from the min to max inclusive using
range(min_val, max_val +1), then convert to a set for easy comparison.Identify Missing Values: The difference between the required range set and the unique integers set gives us the missing values. We sort them to maintain a logical order before adding to the result list.
Edge Case Handling
- Single-Value Ranges: When min and max are the same (like
[1,1]), the code checks if that single integer exists in the column. If not, it will appear in the missing list. - Non-Integer Values: Any non-integer values (strings or non-whole floats) are ignored during conversion, so they don’t interfere with checking for required integers.
- NaNs: These are dropped entirely since they can’t be considered valid integers in your range.
Example Output
Using your provided min-max list [[1,2], [2,4], [3,6], [101,103], [203,206], [2,3], [1,1]], if:
- Column 0 only has
1→ missing[2] - Column 1 has
2,3→ missing[4] - Column 2 has
3,4,6→ missing[5] - Column 3 has all values → missing
[] - Column4 has
203,205,206→ missing[204] - Column5 has
2,3→ missing[] - Column6 has no
1→ missing[1]
The output would be:[[2], [4], [5], [], [204], [], [1]]
内容的提问来源于stack exchange,提问作者user14368706

