You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在保留映射关系的前提下实现Pandas DataFrame列值的互斥化转换

Solution for Mutually Exclusive Categorical Columns with Preserved Type Mappings

Problem Overview

You need to transform your existing categorical mapping dictionary (information_dict_from) and pandas DataFrame (data_from) into a new structure where:

  • Column values are mutually exclusive: No overlap between the value sets of any two columns
  • Original type mappings are preserved: Each original key-to-type relationship is retained, just mapped to a new unique key

Your initial data:

import pandas as pd

information_dict_from = { 
    "v1": {0: "type a", 1: "type b"}, 
    "v2": {0: "type a", 1: "type b", 3: "type c"}, 
    "v3": {0: "type a", 1: "type b"}, 
}
data_from = pd.DataFrame( { 
    "v1": [0, 0, 1, 1], 
    "v2": [0, 1, 1, 3], 
    "v3": [0, 1, 1, 0], 
} )

Desired output:

information_dict_to = { 
    "v1": {0: "type a", 1: "type b"}, 
    "v2": {2: "type a", 3: "type b", 4: "type c"}, 
    "v3": {5: "type a", 6: "type b"}, 
}
data_to = pd.DataFrame( { 
    "v1": [0, 0, 1, 1], 
    "v2": [2, 3, 3, 4], 
    "v3": [5, 6, 6, 5], 
} )

Step-by-Step Solution

The core idea is to use an incremental offset to shift the keys of each column's mapping, ensuring each column's value range starts right after the previous column's maximum value. This guarantees no overlap between columns.

Implementation Code

import pandas as pd

# Initial data
information_dict_from = { 
    "v1": {0: "type a", 1: "type b"}, 
    "v2": {0: "type a", 1: "type b", 3: "type c"}, 
    "v3": {0: "type a", 1: "type b"}, 
}
data_from = pd.DataFrame( { 
    "v1": [0, 0, 1, 1], 
    "v2": [0, 1, 1, 3], 
    "v3": [0, 1, 1, 0], 
} )

# Initialize output structures
information_dict_to = {}
data_to = data_from.copy()
current_offset = 0

# Iterate through each column to update mappings and values
for col in information_dict_from:
    original_mapping = information_dict_from[col]
    # Create new mapping by adding current offset to original keys
    new_mapping = {key + current_offset: value for key, value in original_mapping.items()}
    information_dict_to[col] = new_mapping
    # Update DataFrame column values with the same offset
    data_to[col] = data_to[col] + current_offset
    # Update offset to start after the maximum key of the current column's new mapping
    current_offset = max(new_mapping.keys()) + 1

# Verify the results
print("Generated information_dict_to:")
for k, v in information_dict_to.items():
    print(f"{k}: {v}")

print("\nGenerated data_to:")
print(data_to)

How It Works

  1. Initialization: We start with an offset of 0 and make a copy of the original DataFrame to avoid modifying the input data.
  2. Column-wise Processing:
    • For each column, we create a new mapping by adding the current offset to every original key. This preserves the type association while shifting the key values.
    • We update the corresponding DataFrame column by adding the same offset to all values, maintaining the original categorical logic.
  3. Offset Update: After processing each column, we set the next offset to be one more than the maximum key in the current column's new mapping. This ensures the next column's values start in a completely new range.

Validation

To confirm the mutual exclusivity requirement is met, you can check the value sets of any two columns:

# Check overlap between v1 and v2
print(f"v1 values: {set(data_to['v1'])}")
print(f"v2 values: {set(data_to['v2'])}")
print(f"Overlap between v1 and v2: {set(data_to['v1']) & set(data_to['v2'])}")  # Should be empty set

This will output an empty set, confirming no overlap between the columns.

内容的提问来源于stack exchange,提问作者baxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 19:17:35