如何用Pandas合并含不可哈希字典的列表并去重,规避unhashable type错误?
Great question! Since dictionaries are unhashable, we can't directly use pandas' built-in deduplication functions on them. The workaround is to temporarily serialize these dictionaries into a hashable format (like tuples) for deduplication, then convert them back to dictionaries. Here's a step-by-step implementation:
Step 1: Convert Your List to a DataFrame
First, we'll turn your input list into a pandas DataFrame to leverage grouping functionality:
import pandas as pd # Your input data items = [ { 'FirstName': 'David', 'LastName': 'Smith', 'Residence': [{'Place': 'X', 'Age': 22}, {'Place': 'Y', 'Age': 23}] }, { 'FirstName': 'David', 'LastName': 'Smith', 'Residence': [{'Place': 'Z', 'Age': 20}] }, { 'FirstName': 'David', 'LastName': 'Smith', 'Residence': [{'Place': 'Z', 'Age': 20}] }, { 'FirstName': 'Bob', 'LastName': 'Jones', 'Residence': [{'Place': 'Z', 'Age': 20}] } ] df = pd.DataFrame(items)
Step 2: Define a Deduplication Function
We'll create a function that takes a collection of Residence lists, converts each dictionary to a sorted tuple (to ensure consistent hashing regardless of key order), deduplicates using a set, then converts back to dictionaries:
def merge_deduplicate_residences(residence_groups): unique_res_tuples = set() # Iterate through all Residence lists in the group for res_list in residence_groups: for res_dict in res_list: # Convert dictionary to sorted tuple (hashable and consistent) hashed_res = tuple(sorted(res_dict.items())) unique_res_tuples.add(hashed_res) # Convert tuples back to dictionaries return [dict(res_tuple) for res_tuple in unique_res_tuples]
Step 3: Group and Apply the Function
Use pandas groupby to group by FirstName and LastName, then apply our deduplication function to the Residence column:
# Group by name columns and aggregate Residence result_df = df.groupby(['FirstName', 'LastName'], as_index=False).agg( Residence=('Residence', merge_deduplicate_residences) ) # Convert back to the desired list of dictionaries format final_result = result_df.to_dict('records') print(final_result)
Output
Running this code will produce exactly the result you're looking for:
[ {'FirstName': 'Bob', 'LastName': 'Jones', 'Residence': [{'Place': 'Z', 'Age': 20}]}, {'FirstName': 'David', 'LastName': 'Smith', 'Residence': [{'Place': 'X', 'Age': 22}, {'Place': 'Y', 'Age': 23}, {'Place': 'Z', 'Age': 20}]} ]
Key Notes
- We sort the dictionary items before converting to tuples to avoid treating identical dictionaries with different key orders as unique.
- This method works for flat dictionaries like your example. If you had nested dictionaries, you'd need a recursive serialization function, but that's unnecessary here.
- The
aggmethod lets us customize how we combine grouped data, which is perfect for this "GROUP_CONCAT DISTINCT" style operation.
内容的提问来源于stack exchange,提问作者David542

