You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何以最低时间成本避免创建重复Python类实例?

避免重复创建类实例的最优方案

我正在处理超大数据集,通过循环分块向类中添加元素。数据存在大量重复值,导致多次为相同数据创建类实例。测试显示类实例创建是操作中最耗时的环节,因此需要以最低时间成本避免重复创建实例——同一数据仅创建一次,所有重复项引用该实例。无法预先移除数据中的重复项,需尽量降低耗时操作。

以下是用于说明问题的示例代码(原代码运行时间约4.22秒,预期优化后约2.8秒):

import time
from collections import defaultdict

SLEEP_1 = 0.2
SLEEP_2 = 0.5

# 模拟实例创建有较高时间成本的Person类
class Person:
    def __init__(self, info):
        self._id = info['_id']
        self.name = info['name']
        self.nationality = info['nationality']
        self.age = info['age']
        self.can_drink_in_USA = self.some_long_fun()
        self.can_fly_solo = self.another_costly_fun()

    def some_long_fun(self):
        time.sleep(SLEEP_1)
        return self.age >= 21

    def another_costly_fun(self):
        time.sleep(SLEEP_2)
        return self.age >= 18


# 包含重复数据的测试数据集(James出现3次)
teams = {
    "team1": [
        {"_id": "foo", "name": "James", "nationality": "French", "age": 32},
        {"_id": "bar", "name": "Frank", "nationality": "American", "age": 36},
        {"_id": "foo", "name": "James", "nationality": "French", "age": 32}
    ],
    "team2": [
        {"_id": "foo", "name": "James", "nationality": "French", "age": 32},
        {"_id": "baz", "name": "Oliver", "nationality": "British", "age": 26},
        {"_id": "qux", "name": "Josh", "nationality": "British", "age": 42}
    ]
}

优化方案

核心思路是用缓存字典存储已创建的实例,以数据的唯一标识(如_id)作为键。每次处理数据时先检查缓存:存在则直接引用已有的实例,不存在则创建新实例并存入缓存。这种方式的时间成本极低,仅需一次字典查找(O(1)复杂度)。

修改后的完整代码:

import time
from collections import defaultdict

SLEEP_1 = 0.2
SLEEP_2 = 0.5

class Person:
    def __init__(self, info):
        self._id = info['_id']
        self.name = info['name']
        self.nationality = info['nationality']
        self.age = info['age']
        self.can_drink_in_USA = self.some_long_fun()
        self.can_fly_solo = self.another_costly_fun()

    def some_long_fun(self):
        time.sleep(SLEEP_1)
        return self.age >= 21

    def another_costly_fun(self):
        time.sleep(SLEEP_2)
        return self.age >= 18


teams = {
    "team1": [
        {"_id": "foo", "name": "James", "nationality": "French", "age": 32},
        {"_id": "bar", "name": "Frank", "nationality": "American", "age": 36},
        {"_id": "foo", "name": "James", "nationality": "French", "age": 32}
    ],
    "team2": [
        {"_id": "foo", "name": "James", "nationality": "French", "age": 32},
        {"_id": "baz", "name": "Oliver", "nationality": "British", "age": 26},
        {"_id": "qux", "name": "Josh", "nationality": "British", "age": 42}
    ]
}


person_cache = {}  # 用字典缓存已创建的Person实例
team_directory = defaultdict(list)

start_time = time.time()
for team_name, members in teams.items():
    for idx, person_info in enumerate(members):
        person_id = person_info['_id']
        if person_id in person_cache:
            print(f"{person_info['name']} [_id: {person_id}] 已存在,直接引用实例")
            p = person_cache[person_id]
        else:
            print(f"创建新实例:Person {idx + 1} = {person_info['name']}")
            p = Person(info=person_info)
            person_cache[person_id] = p
        team_directory[team_name].append(p)

finish_time = time.time() - start_time
expected_finish = round((SLEEP_1 * 4) + (SLEEP_2 * 4), 2)
print(f"构建团队目录耗时:{round(finish_time, 2)}s [预期:{expected_finish}s]")

# 验证结果:每个团队保留3个成员,重复项引用同一实例
for team_name, members in team_directory.items():
    roster = " ".join([p.name for p in members])
    print(f"Team {team_name} 成员:{roster}")
    # 验证重复实例是否为同一对象
    if team_name == "team1":
        print(f"James实例是否为同一对象:{members[0] is members[2]}")  # 输出True

优化效果

  • 仅为唯一的4个_id创建实例,避免了重复创建James的实例
  • 实际运行时间接近预期的2.8秒,大幅降低了耗时
  • 所有重复数据项均引用同一实例,不影响后续业务逻辑

内容的提问来源于stack exchange,提问作者fugu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 15:40:36