如何在保留映射关系的前提下实现Pandas DataFrame列值的互斥化转换
Solution for Mutually Exclusive Categorical Columns with Preserved Type Mappings
Problem Overview
You need to transform your existing categorical mapping dictionary (information_dict_from) and pandas DataFrame (data_from) into a new structure where:
- Column values are mutually exclusive: No overlap between the value sets of any two columns
- Original type mappings are preserved: Each original key-to-type relationship is retained, just mapped to a new unique key
Your initial data:
import pandas as pd information_dict_from = { "v1": {0: "type a", 1: "type b"}, "v2": {0: "type a", 1: "type b", 3: "type c"}, "v3": {0: "type a", 1: "type b"}, } data_from = pd.DataFrame( { "v1": [0, 0, 1, 1], "v2": [0, 1, 1, 3], "v3": [0, 1, 1, 0], } )
Desired output:
information_dict_to = { "v1": {0: "type a", 1: "type b"}, "v2": {2: "type a", 3: "type b", 4: "type c"}, "v3": {5: "type a", 6: "type b"}, } data_to = pd.DataFrame( { "v1": [0, 0, 1, 1], "v2": [2, 3, 3, 4], "v3": [5, 6, 6, 5], } )
Step-by-Step Solution
The core idea is to use an incremental offset to shift the keys of each column's mapping, ensuring each column's value range starts right after the previous column's maximum value. This guarantees no overlap between columns.
Implementation Code
import pandas as pd # Initial data information_dict_from = { "v1": {0: "type a", 1: "type b"}, "v2": {0: "type a", 1: "type b", 3: "type c"}, "v3": {0: "type a", 1: "type b"}, } data_from = pd.DataFrame( { "v1": [0, 0, 1, 1], "v2": [0, 1, 1, 3], "v3": [0, 1, 1, 0], } ) # Initialize output structures information_dict_to = {} data_to = data_from.copy() current_offset = 0 # Iterate through each column to update mappings and values for col in information_dict_from: original_mapping = information_dict_from[col] # Create new mapping by adding current offset to original keys new_mapping = {key + current_offset: value for key, value in original_mapping.items()} information_dict_to[col] = new_mapping # Update DataFrame column values with the same offset data_to[col] = data_to[col] + current_offset # Update offset to start after the maximum key of the current column's new mapping current_offset = max(new_mapping.keys()) + 1 # Verify the results print("Generated information_dict_to:") for k, v in information_dict_to.items(): print(f"{k}: {v}") print("\nGenerated data_to:") print(data_to)
How It Works
- Initialization: We start with an offset of 0 and make a copy of the original DataFrame to avoid modifying the input data.
- Column-wise Processing:
- For each column, we create a new mapping by adding the current offset to every original key. This preserves the type association while shifting the key values.
- We update the corresponding DataFrame column by adding the same offset to all values, maintaining the original categorical logic.
- Offset Update: After processing each column, we set the next offset to be one more than the maximum key in the current column's new mapping. This ensures the next column's values start in a completely new range.
Validation
To confirm the mutual exclusivity requirement is met, you can check the value sets of any two columns:
# Check overlap between v1 and v2 print(f"v1 values: {set(data_to['v1'])}") print(f"v2 values: {set(data_to['v2'])}") print(f"Overlap between v1 and v2: {set(data_to['v1']) & set(data_to['v2'])}") # Should be empty set
This will output an empty set, confirming no overlap between the columns.
内容的提问来源于stack exchange,提问作者baxx
相关产品推荐
相关产品推荐

