如何检测DataFrame中包含列表或字典类型的列并存储结果
First, let's start by recreating your sample DataFrame to work with:
import pandas as pd # Your sample DataFrame df = pd.DataFrame({ 'col1': ['A', 'B', 'C'], 'col2': [11, 21, 31], 'col3': [[{'id':2}], [{'id':3}], [{'id':4}]], 'col4': [{"price": 0.0}, {"price": 2.0}, {"price": 3.0}] })
Approach 1: Robust Check for Any List/Dict Elements (Handles Mixed Types)
This method checks all elements in each column to see if any are lists or dictionaries, making it suitable even if columns have mixed data types:
# Get unique data types for each column column_unique_types = df.apply(lambda col: col.apply(type).unique()) # Filter columns containing list or dict types list_dict_columns = [ col_name for col_name, types in column_unique_types.items() if any(dt in (list, dict) for dt in types) ] print(list_dict_columns) # Output: ['col3', 'col4']
How it works:
df.apply(lambda col: col.apply(type).unique()): For each column, we apply thetype()function to every element, then get the unique types present in that column.- We iterate over each column's unique types. If any type is
listordict, we add the column name to our result list.
Approach 2: Simplified Check (Assumes Uniform Column Types)
If you're confident all elements in a column are the same type (common in most cases), you can check just the first element for efficiency:
list_dict_columns = [] for col in df.columns: first_element_type = type(df[col].iloc[0]) if first_element_type in (list, dict): list_dict_columns.append(col) print(list_dict_columns) # Output: ['col3', 'col4']
Approach 3: Using Value Counts (Building on Your Initial Attempt)
If you want to build on your original applymap(type).apply(pd.value_counts) approach, here's how to use it to get the desired columns:
# Get type counts per column type_counts = df.applymap(type).apply(pd.value_counts) # Filter columns where list or dict is present in the type counts list_dict_columns = [ col for col in type_counts.columns if any(dt in (list, dict) for dt in type_counts[col].dropna().index) ] print(list_dict_columns) # Output: ['col3', 'col4']
This works by checking if list or dict exists in the index of the value counts (after dropping NaN values for types that don't appear in the column).
All three methods will give you the list of columns containing list or dictionary elements. Choose the one that best fits your data's structure!
内容的提问来源于stack exchange,提问作者Shabari nath k

