使用Numpy统计CSV唯一值遇类型错误,调用masked_invalid仍报错
Got it, let's break down why you're hitting these errors and fix them step by step.
First, the root cause here is that your CSV matrix has a mix of string values and NaN (which is a float type in NumPy). When you try np.unique(), it attempts to sort elements to identify unique values—and you can't compare strings and floats directly, hence the TypeError: '<' not supported between instances of 'str' and 'float' error.
As for np.ma.masked_invalid(), that function is built for numeric arrays: it checks if elements are finite numbers. Since your array contains strings, the isfinite ufunc can't handle those types, leading to the second error you saw.
Here are two solid, pure NumPy solutions tailored to your needs:
Solution 1: Filter out NaN values first, then get unique values
If you don't need to include NaN in your unique count, create a mask to keep only valid non-NaN elements before running np.unique():
import numpy as np # Assume csv_matrix is your mixed-type 2D array (strings + NaN floats) def is_non_nan(x): # Check if the element is a float AND NaN; keep everything else return not (isinstance(x, float) and np.isnan(x)) # Apply the check across the entire array to create a boolean mask valid_mask = np.vectorize(is_non_nan)(csv_matrix) # Extract valid elements and flatten to a 1D array for np.unique valid_elements = csv_matrix[valid_mask].flatten() # Get your unique values unique_values = np.unique(valid_elements)
Solution 2: Convert NaN to a string placeholder, then get unique values
If you want to count NaN as a distinct unique value, convert all NaN floats to a consistent string placeholder (like 'NaN') so the entire array uses the same type:
import numpy as np # Flatten the matrix, replace NaN with 'NaN' string, then reshape back cleaned_matrix = np.array([ x if not (isinstance(x, float) and np.isnan(x)) else 'NaN' for x in csv_matrix.flatten() ]).reshape(csv_matrix.shape) # Now np.unique works seamlessly since all elements are strings unique_values = np.unique(cleaned_matrix)
Either approach will resolve the type mismatch issues you're facing. Pick the one that aligns with whether you want to exclude NaN or treat it as a unique entry.
内容的提问来源于stack exchange,提问作者yololala

