Python函数优化咨询:精简assign_points,避免重复遍历与全局变量
问题描述
输入数据示例
[['0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0'], ['0', '0', '10', '10', '0', '0', '10', '10', '0', '0', '0', '10', '10', '10', '10', '10', '0', '0', '0'], ['0', '0', '10', '10', '0', '0', '10', '10', '0', '0', '0', '0', '0', '0', '0', '10', '10', '0', '0'], ['0', '0', '10', '10', '0', '0', '10', '10', '0', '0', '0', '0', '0', '0', '0', '10', '10', '0', '0'], ['0', '0', '10', '10', '10', '10', '10', '10', '0', '0', '0', '0', '10', '10', '10', '10', '0', '0', '0'], ['0', '0', '0', '10', '10', '10', '10', '10', '0', '0', '0', '10', '10', '0', '0', '0', '0', '0', '0'], ['0', '0', '0', '0', '0', '0', '10', '10', '0', '0', '0', '10', '10', '0', '0', '0', '0', '0', '0'], ['0', '0', '0', '0', '0', '0', '10', '10', '0', '0', '0', '10', '10', '10', '10', '10', '10', '0', '0'], ['0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0'], ['0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0', '0']]
当前实现函数
def assign_points(lists): """ Assigns points from a list of lists to Vec4 objects. ... """ z_max = float('-inf') z_min = float('inf') points = [] for y, l in enumerate(lists): for x, point in enumerate(l): try: z, color = parse_point(point) if color is None: color = no_color(z) points.append(Vec4(x, y, z, 0, color)) # Update z_max and z_min if necessary if z > z_max: z_max = z if z < z_min: z_min = z except MyWarning as e: exit(f"Error on line {str(y)}, item {str(x)}: {e}") return points, z_max, z_min
核心疑问
assign_points函数职责混杂,代码显得臃肿- 输入列表规模可能很大,担心多次遍历会损耗性能
- 不确定全局变量是否是最优选择
- 考虑过用带标志参数的函数存储状态并返回最值,但不清楚Python是否支持静态变量,纠结是维持现状、用全局变量还是其他方案
解决方案
1. 拆分职责,保持一次遍历的性能优势
不要为了拆分而增加遍历次数,而是把单个点的处理逻辑抽离成辅助函数,让主函数只负责统筹遍历和状态更新,既简化代码又不损失性能:
def _process_single_point(x, y, point_str): """处理单个点字符串,返回Vec4对象和对应的z值""" z, color = parse_point(point_str) if color is None: color = no_color(z) return Vec4(x, y, z, 0, color), z def assign_points(lists): """ Assigns points from a list of lists to Vec4 objects, and calculates z range. ... """ z_max = float('-inf') z_min = float('inf') points = [] for y, row in enumerate(lists): for x, point_str in enumerate(row): try: vec4_point, z = _process_single_point(x, y, point_str) points.append(vec4_point) # 更新z的最值 if z > z_max: z_max = z if z < z_min: z_min = z except MyWarning as e: exit(f"Error on line {str(y)}, item {str(x)}: {e}") return points, z_max, z_min
2. 替代全局变量/静态变量的方案
Python没有直接的静态变量,且全局变量会提升代码耦合度,不利于维护和测试,不推荐使用,可以用以下两种更合理的方式:
方案A:用类封装状态
把点列表、z最值等状态封装在类实例中,逻辑更内聚,也方便后续扩展:
class PointProcessor: def __init__(self): self.z_max = float('-inf') self.z_min = float('inf') self.points = [] def process(self, lists): for y, row in enumerate(lists): for x, point_str in enumerate(row): try: z, color = parse_point(point_str) if color is None: color = no_color(z) self.points.append(Vec4(x, y, z, 0, color)) if z > self.z_max: self.z_max = z if z < self.z_min: self.z_min = z except MyWarning as e: exit(f"Error on line {str(y)}, item {str(x)}: {e}") return self.points, self.z_max, self.z_min # 使用方式 processor = PointProcessor() points, z_max, z_min = processor.process(lists)
方案B:用可变对象传递状态
如果不想用类,可以用字典这类可变对象来承载最值状态,在遍历中更新:
def update_z_range(z, z_range): """将单个z值传入字典,更新其中的最值""" if z > z_range['max']: z_range['max'] = z if z < z_range['min']: z_range['min'] = z def assign_points(lists): z_range = {'max': float('-inf'), 'min': float('inf')} points = [] for y, row in enumerate(lists): for x, point_str in enumerate(row): try: vec4_point, z = _process_single_point(x, y, point_str) points.append(vec4_point) update_z_range(z, z_range) except MyWarning as e: exit(f"Error on line {str(y)}, item {str(x)}: {e}") return points, z_range['max'], z_range['min']
3. 性能总结
当前一次遍历的逻辑已经是最优的,时间复杂度为O(n*m)(n为行数,m为列数)。无论怎么拆分,只要保持一次遍历完成点生成和最值计算,就不会有性能浪费;如果拆分后进行两次遍历,时间复杂度不变,但实际运行效率会有所下降。
最终建议
- 放弃全局变量,避免代码耦合
- 优先选择拆分辅助函数+一次遍历的方案,兼顾代码清晰度和性能
- 如果需要更内聚的逻辑,用类封装状态是比全局变量、静态变量更符合Python编码习惯的选择
内容的提问来源于stack exchange,提问作者MiguelP
相关产品推荐
相关产品推荐

