You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中是否存在更优的字符串双向映射(可逆字典)实现方案?

存在现成的「可逆字符串-整数映射」工具吗?

你描述的这种“字符串与整数双向映射、用索引节省空间”的需求,Python生态里有不少成熟工具,完全不用自己从零实现:

1. scikit-learn 的 LabelEncoder

这是机器学习场景里最常用的工具,专门做类别(字符串/离散值)到整数的双向映射,完全匹配你的需求:

  • 支持批量拟合唯一值,自动生成连续索引
  • 提供transform(字符串转整数)和inverse_transform(整数转字符串)方法
  • 内部实现逻辑和你的StrMap一致,但性能经过官方优化,稳定性更强

示例用法:

from sklearn.preprocessing import LabelEncoder

le = LabelEncoder()
# 拟合所有唯一值
le.fit(["北京", "上海", "广州", "北京"])
# 字符串转整数
print(le.transform(["北京", "上海"]))  # 输出 [0, 1]
# 整数转字符串
print(le.inverse_transform([0, 2]))  # 输出 ['北京', '广州']

2. pandas 的 Categorical

如果你在处理大规模数据集(比如你提到的百万条数据库记录),pandas的Categorical类型天生就是为这种场景设计的:

  • 把字符串列转成类别类型后,自动用整数索引存储,能大幅降低内存占用
  • 通过cat.codes获取整数索引,cat.categories获取原字符串列表
  • 直接兼容DataFrame/Series操作,不用额外写适配代码

示例用法:

import pandas as pd

s = pd.Series(["北京", "上海", "广州", "北京"])
cat_s = s.astype("category")
# 查看整数索引
print(cat_s.cat.codes)  # 输出 0,1,2,0
# 查看原字符串类别
print(cat_s.cat.categories)  # 输出 Index(['北京', '上海', '广州'], dtype='object')
# 整数转字符串
print(cat_s.cat.categories[0])  # 输出 '北京'

3. bidict 库

这是一个专门的双向字典库,支持键值对的双向快速查找,虽然不像前两个工具针对“字符串转连续索引”做优化,但胜在通用性强:

  • 可以直接通过bidict_obj[key]正向查找、bidict_obj.inverse[value]反向查找
  • 支持动态添加键值对,维护成熟,文档完善

示例用法:

from bidict import bidict

bd = bidict({"北京":0, "上海":1})
print(bd["北京"])  # 输出 0
print(bd.inverse[1])  # 输出 "上海"

和你实现的StrMap对比

你自己写的StrMap思路完全正确,但这些现成工具的优势在于:

  • 经过大量用户验证,bug更少
  • 支持更多边缘场景(比如处理空值、批量操作、异常捕获)
  • 和现有数据处理/机器学习生态兼容,不用自己再写适配逻辑

你的实现代码参考:

class StrMap(dict):
    """ provides a fast mapping of strings to integer indices in a list.
        all keys of type int have values of type str
        all keys of type str have values of type int
        examples: d['ace'] : 0,  d['bob'] : 1, d[0] = 'ace', d[1] = 'bob' and so on
    """    
    def __init__(self):
        dict.__init__(self)
        self.L = []

    def add(self, newval : str) -> int:
        assert isinstance(newval, str)
        if newval not in self:
            newidx = len(self.L)
            self.L.append(newval)
            self[newval] = newidx
            return newidx

    def retr(self, multikey, addOnMiss=True):
        if isinstance(multikey, int):
            return self.L[multikey]
        elif isinstance(multikey, str):
            try:
                return self[multikey]
            except KeyError:
                if addOnMiss:
                    return self.add(multikey)
                else:
                    print(f"addOnMiss is not set and string key {multikey} is not in dict")
                    raise KeyError

内容的提问来源于stack exchange,提问作者ljwobker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 00:19:59