Python循环实例化对象会引发内存泄漏吗?多模拟耗时递增排查
我用Python 3.9的面向对象编程实现了一个基于主体的捕食者-猎物种群动态模型,模拟变化环境中的种群行为。当通过for循环运行多轮模拟时,单轮模拟的运行时间随轮次增加而变长,怀疑存在内存泄漏但无法定位原因。
代码框架
# Parameters n_deers = ... n_wolves = ... # etc. # Functions def some_function(arg): pass # Helper objects some_dict = ... # Classes class Deer: pass class Wolf: pass class Environment: def __init__(self): self.deers = [Deer(ID = i) for i in range(n_deers)] self.wolves = [Wolf(ID = i) for i in range(n_wolves)] self.data = pd.DataFrame() def simulation(self): pass # Simulations for i in range(100): environment = Environment() environment.simulation() environment.data.to_csv()
我定义了全局参数、全局函数和供类实例调用的全局字典;为每种动物定义了类,还定义了Environment类,其内部会生成一定数量的动物实例,并在模拟运行中通过DataFrame追踪动物的移动、觅食、繁殖、死亡等行为。我担心每次模拟的动物实例(完整模拟约7000只)未被正确回收,但并未使用静态类变量。
后续定位到问题根源在动物移动模块,以下是最小可复现代码:
最小可复现代码
# MINIMAL MOVEMENT MODEL # IMPORTS import random as rd import numpy as np import time import psutil # REPRODUCIBILITY rd.seed(42) # PARAMETERS landscape_size = 11 n_deers = 100 years = 10 length_year = 360 timesteps = years*length_year n_simulations = 20 # HELPER FUNCTIONS AND OBJECTS # Landscape for first initialization mock_landscape = np.zeros((landscape_size,landscape_size)) # Function to return a list of nxn cells around a given cell def range_finder(matrix, position, radius): adj = [] lower = 0 - radius upper = 1 + radius for dx in range(lower, upper): for dy in range(lower, upper): rangeX = range(0, matrix.shape[0]) # Identifies X bounds rangeY = range(0, matrix.shape[1]) # Identifies Y bounds (newX, newY) = (position[0]+dx, position[1]+dy) # Identifies adjacent cell if (newX in rangeX) and (newY in rangeY) and (dx, dy) != (0, 0): adj.append((newX, newY)) return adj # Nested dictionary that contains all sets of neighbors for all possible distances up to half the landscape size neighbor_dict = {d: {(i,j): range_finder(mock_landscape, (i,j), d) for i in range(landscape_size) for j in range(landscape_size)} for d in range(1,int(landscape_size/2)+1)} # Function that picks the cell in the home range that was visited longest ago def cell_choice(position, home_range, memory): # These are all the adjacent cells to the current position adjacent_cells = neighbor_dict[1][position] # This is the subset of cells of the adjacent cells belonging to homerange possible_choices = [i for i in adjacent_cells if i in home_range] # This yields the "master" indeces of those choices indeces = [] for i in possible_choices: indeces.append(home_range.index(i)) # This picks the index with the maximum value in the memory (ie visited longest ago) memory_values = [memory[i] for i in indeces] pick_index = indeces[memory_values.index(max(memory_values))] # Sets that values memory to zero memory[pick_index] = 0 # # Adds one period to every other index other_indeces = [i for i in list(range(len(memory))) if i != pick_index] for i in other_indeces: memory[i] += 1 # Returns the picked cell return home_range[pick_index] # CLASS DEFINITIONS class Deer: def __init__(self, ID): self.ID = ID self.position = (rd.randint(0,landscape_size-1),rd.randint(0,landscape_size-1)) # Sets up a counter how long the deer has been in the cell self.time_spent_in_cell = 1 # Defines a distance parameter that specifies the radius of the homerange around the base self.movement_radius = 1 # Defines an initial home range around the position self.home_range = neighbor_dict[self.movement_radius][self.position] self.home_range.append(self.position) # Sets up a list of counters how long ago cells in the home range have been visited self.memory = [float('inf')]*len(self.home_range) self.memory[self.home_range.index(self.position)] = 0 def move(self): self.position = cell_choice(self.position, self.home_range, self.memory) class Environment: def __init__(self): self.landscape = np.zeros((landscape_size, landscape_size)) self.deers = [Deer(ID = i) for i in range(n_deers)] def simulation(self): for timestep in range(timesteps): for deer in self.deers: deer.move() # SIMULATIONS process = psutil.Process() times = [] memory = [] for i in range(1,n_simulations+1): print(i, " out of ",n_simulations) start_time = time.time() environment = Environment() environment.simulation() times.append(time.time() - start_time) memory.append(process.memory_info().rss) print(times) print(memory)
说明:当让动物随机选择相邻单元格移动时,问题未出现;但添加记忆、家域和cell_choice()函数后,模拟耗时随轮次递增——在我的机器上,第一轮模拟耗时3-4秒,最后一轮耗时10-11秒。
核心原因
问题不在于内存泄漏,而是**cell_choice函数中使用的float('inf')在循环累加过程中会逐渐引发性能损耗**:
- Deer初始化时,
memory列表除当前位置外都被设为float('inf') - 每次调用
cell_choice时,未被选中的memory元素会执行memory[i] += 1操作 float('inf') + 1仍然是inf,但当某个元素被选中重置为0后,后续累加会变成普通整数;随着模拟轮次增加,越来越多的元素从inf转为整数并持续累加,这些不断增大的整数会占用更多计算资源,导致每轮模拟的耗时逐步上升。
修复方案
将memory初始化的float('inf')替换为一个足够大的初始整数(比如timesteps的2倍),避免使用无限大值:
# Deer类__init__中原代码 self.memory = [float('inf')]*len(self.home_range) # 修改为 self.memory = [2 * timesteps] * len(self.home_range)
额外优化建议
将
home_range从列表转为集合:原代码中possible_choices = [i for i in adjacent_cells if i in home_range]使用列表的in操作是O(n)复杂度,转为集合后in操作是O(1),能大幅提升性能:# Deer类__init__中修改 self.home_range = set(neighbor_dict[self.movement_radius][self.position]) self.home_range.add(self.position)预计算位置到索引的映射:原代码中
home_range.index(i)是O(n)操作,可在Deer初始化时预计算映射字典,避免重复查找:# Deer类__init__中添加 self.pos_to_idx = {pos: idx for idx, pos in enumerate(list(self.home_range))} # cell_choice中替换indeces的生成 indeces = [self.pos_to_idx[i] for i in possible_choices]使用numpy数组存储memory:numpy处理数值运算比纯Python列表更快,可进一步降低计算开销:
# Deer类__init__中修改 self.memory = np.full(len(self.home_range), 2 * timesteps, dtype=np.int32) self.memory[self.pos_to_idx[self.position]] = 0 # cell_choice中调整相关操作 memory_values = self.memory[indeces] pick_index = indeces[np.argmax(memory_values)] self.memory[pick_index] = 0 self.memory += 1 self.memory[pick_index] = 0 # 抵消全局加1的操作,保持选中位置为0
内容的提问来源于stack exchange,提问作者Peter Kamal

