Pandas列布尔数组索引报错求助:筛选f500数据集数值列
Hey there! Let's sort out that column indexing error you're hitting with your f500 DataFrame. The issue comes down to how pandas handles column-based boolean indexing vs. numpy-style array slicing.
What went wrong
Your line f500[:, numeric_only_bool] uses numpy's array slicing syntax, but pandas doesn't interpret [:, ...] the same way for DataFrames. The first position in [] is reserved for row selection, so passing a boolean array there (meant for columns) causes an error.
Fix 1: Use .loc for explicit row/column selection
Pandas' .loc method lets you explicitly specify row and column selectors. Since you want all rows and only numeric columns, you can do:
# Your original boolean mask is correct numeric_only_bool = f500.dtypes != object # Fix the indexing with .loc numeric_only = f500.loc[:, numeric_only_bool]
The : in the first position means "select all rows", and the boolean array in the second position filters which columns to keep.
Fix 2: Use select_dtypes (even simpler!)
Pandas has a built-in method specifically for filtering columns by data type, which is more concise and reliable than manual boolean masks. To get all numeric columns:
numeric_only = f500.select_dtypes(include=['number'])
If you want to narrow it down to only integers and floats (exclude nullable numeric types like Int64), you can specify the exact dtypes:
numeric_only = f500.select_dtypes(include=['int64', 'float64'])
Either approach will give you the subset of your DataFrame with only numeric columns. Let me know if you run into any other issues!
内容的提问来源于stack exchange,提问作者jaeseongpark

