NumPy是否有类似Pandas的select_dtypes方法?实现按类型筛选
select_dtypes NumPy doesn’t have a built-in select_dtypes function like Pandas, but we can easily replicate this functionality using structured arrays—NumPy’s way of handling tabular data with mixed data types. Here’s a step-by-step implementation:
Step 1: Convert your Pandas DataFrame to a NumPy structured array
Standard NumPy arrays are homogeneous (all elements share the same dtype), so we first convert the DataFrame to a structured array (which supports mixed dtypes) using df.to_records():
import pandas as pd import numpy as np # Create your sample DataFrame df = pd.DataFrame({'a': [1, 2] * 3, 'b': [True, False] * 3, 'c': [1.0, 2.0] * 3}) # Convert to structured array (exclude the index for cleaner output) np_structured_arr = df.to_records(index=False)
Step 2: Build the dtype-filtering function
We’ll create a custom function that filters fields in the structured array based on their dtypes, mirroring Pandas’ select_dtypes behavior:
def select_dtypes_numpy(arr, include=None, exclude=None): # Set default empty lists if parameters aren't provided include = include or [] exclude = exclude or [] # Helper to convert dtype strings (like 'float64') to actual NumPy dtype objects def resolve_dtype(dt): return np.dtype(dt) if isinstance(dt, str) else dt # Normalize include/exclude lists to use consistent dtype objects include_dtypes = [resolve_dtype(dt) for dt in include] exclude_dtypes = [resolve_dtype(dt) for dt in exclude] # Get all field names and their corresponding dtypes field_details = [(name, arr.dtype[name]) for name in arr.dtype.names] # Filter fields that match our rules selected_fields = [] for field_name, field_dtype in field_details: # Skip if the field is explicitly excluded if exclude_dtypes and field_dtype in exclude_dtypes: continue # Keep the field if it's included (or if include list is empty) if not include_dtypes or field_dtype in include_dtypes: selected_fields.append(field_name) # Return the filtered structured array return arr[selected_fields]
Step 3: Use the function to replicate your Pandas example
Now we can use this function to select only float64 fields, just like your Pandas code:
# Select float64 fields float_only_arr = select_dtypes_numpy(np_structured_arr, include=['float64']) print(float_only_arr)
Output:
[(1.,) (2.,) (1.,) (2.,) (1.,) (2.,)]
If you want to convert this back to a Pandas DataFrame for readability, simply run:
pd.DataFrame(float_only_arr)
This will give you the exact same output as df.select_dtypes(include=['float64']).
Extra features
- Use the
excludeparameter to omit specific dtypes, e.g.,select_dtypes_numpy(arr, exclude=['bool'])would remove the boolean field from your sample data. - The function works with both dtype strings (like
'int64') and NumPy dtype objects (likenp.int64).
内容的提问来源于stack exchange,提问作者Student

