如何降低字典的内存占用?附快速属性查找场景需求
嘿,这个场景我太熟悉了——当你要创建大量Plane实例时,字典的内存开销确实会悄悄拖慢你的应用,尤其是每个properties里还塞了一堆小字典。咱们来聊聊几个实用的优化方向,一步步降低内存占用:
1. 合并零散的小字典(最立竿见影的第一步)
看你的例子里,properties是一个字典列表,每个字典居然只有一个键值对?比如[{"canFly": True}, {"isWaterProof": False}]——这完全是在浪费内存啊!每个字典本身都有额外的开销(比如哈希表结构、元数据存储),一个小字典的开销可能比它存的键值对还大。
直接把这些小字典合并成一个大字典:
class Plane(object): def __init__(self, name, properties): self.name = name self.properties = properties # 现在是单个字典,不是列表 @classmethod def from_idx(cls, idx): if idx == 0: return cls("PaperPlane", {"canFly": True, "isWaterProof": False}) if idx == 1: return cls("AirbusA380", {"canFly": True, "maxSpeed": 900, ...})
这样一来,你就把多个字典的开销压缩成了一个,内存占用直接砍半(甚至更多)。
2. 共享重复的字典实例
如果多个Plane实例有相同的属性(比如大部分飞机都有{"canFly": True}),别每次都新建一个字典!把这些重复的字典抽成全局常量,让所有需要的实例共享同一个对象:
# 定义全局共享的属性字典,只创建一次 COMMON_PROPS = { "CAN_FLY_TRUE": {"canFly": True}, "WATER_PROOF_FALSE": {"isWaterProof": False}, "MAX_SPEED_900": {"maxSpeed": 900} } class Plane(object): def __init__(self, name, properties): self.name = name self.properties = properties @classmethod def from_idx(cls, idx): if idx == 0: # 直接引用共享字典,不再新建 return cls("PaperPlane", {**COMMON_PROPS["CAN_FLY_TRUE"], **COMMON_PROPS["WATER_PROOF_FALSE"]}) if idx == 1: return cls("AirbusA380", {**COMMON_PROPS["CAN_FLY_TRUE"], **COMMON_PROPS["MAX_SPEED_900"], ...})
Python里对象是引用传递的,所以这些共享字典只会在内存中存在一次,不管多少个Plane引用它。注意如果你的属性是只读的,这个方法完全安全;如果需要修改属性,那得小心别影响其他实例(可以在需要修改时再复制一份)。
3. 用轻量级结构代替字典:namedtuple或dataclass
字典的灵活性代价就是内存开销,如果你能提前确定属性的键名,用namedtuple或者冻结的dataclass替代字典会更高效——它们没有字典的哈希表 overhead,存储更紧凑。
用namedtuple(Python 2/3都支持)
from collections import namedtuple # 定义固定结构的属性类型,键名提前确定 PlaneProps = namedtuple("PlaneProps", ["canFly", "isWaterProof", "maxSpeed"]) # 创建共享的属性实例,只初始化一次 PAPER_PLANE_PROPS = PlaneProps(canFly=True, isWaterProof=False, maxSpeed=None) AIRBUS_PROPS = PlaneProps(canFly=True, isWaterProof=True, maxSpeed=900) class Plane(object): def __init__(self, name, properties): self.name = name self.properties = properties @classmethod def from_idx(cls, idx): if idx == 0: return cls("PaperPlane", PAPER_PLANE_PROPS) if idx == 1: return cls("AirbusA380", AIRBUS_PROPS)
用dataclass(Python 3.7+,更灵活)
from dataclasses import dataclass # 冻结的dataclass不可变,内存占用和namedtuple差不多,但语法更友好 @dataclass(frozen=True) class PlaneProps: canFly: bool isWaterProof: bool maxSpeed: int | None = None # 同样创建共享实例 PAPER_PLANE_PROPS = PlaneProps(canFly=True, isWaterProof=False) AIRBUS_PROPS = PlaneProps(canFly=True, isWaterProof=True, maxSpeed=900) # Plane类用法和上面一致
这两种结构的内存占用比字典小30%-50%左右,而且访问属性时可以用.语法(比如plane.properties.canFly),比字典的[]更直观。
4. 给Plane类加上__slots__优化实例内存
最后,如果你要创建成千上万的Plane实例,可以给Plane类加上__slots__,让Python不用字典存储实例的属性,而是用固定的数组,进一步减少每个实例的内存开销:
class Plane(object): __slots__ = ["name", "properties"] # 限制实例只能有这两个属性 def __init__(self, name, properties): self.name = name self.properties = properties # from_idx方法不变
这个优化针对的是Plane实例本身,虽然不是直接优化字典,但当实例数量很大时,效果非常明显。
总结一下优先级:先合并零散小字典 → 再共享重复属性实例 → 然后用namedtuple/dataclass替代字典 → 最后加__slots__优化实例。根据你的实际场景选最合适的组合就行!
内容的提问来源于stack exchange,提问作者Merlin1896

