Python迭代器分块遇最大递归深度错误,求可行解决方案
问题原因与解决方案
错误原因
原代码中set_gen()方法使用islice(self, NUM_CHUNKS)迭代自身实例,而迭代self会触发__next__方法,__next__调用stream_dict,当curr_gen耗尽时又会重新调用set_gen(),形成无限递归循环,最终触发最大递归深度错误。
修改后代码
import subprocess from itertools import islice import json NUM_CHUNKS = 10_000 class Sub: def __init__(self, cmd: str): self.p = subprocess.Popen( cmd, shell=True, executable="/bin/bash", stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True # 直接读取文本行,避免手动解码 ) self.curr_gen = None self._next = self.stream_dict def set_gen(self): # 直接从subprocess的stdout读取行,避免迭代自身导致递归 while True: # 一次性读取NUM_CHUNKS行 chunk_lines = list(islice(self.p.stdout, NUM_CHUNKS)) if not chunk_lines: # 子进程输出耗尽,检查子进程状态并退出 self.p.wait() return # 处理每行,去除可能的换行符和空行 cleaned_lines = [line.rstrip('\n') for line in chunk_lines if line.strip()] if not cleaned_lines: continue # 拼接成JSON数组批量解析 data = "[" + ",".join(cleaned_lines) + "]" try: dict_list = json.loads(data) for item in dict_list: yield item except json.JSONDecodeError as e: # 可根据需要添加错误处理逻辑,比如跳过错误行或记录日志 print(f"解析JSON块出错: {e}") continue def stream_dict(self): try: return next(self.curr_gen) except StopIteration: # curr_gen耗尽时,重新生成新的生成器 self.curr_gen = self.set_gen() try: return next(self.curr_gen) except StopIteration: # 所有数据处理完成,触发迭代结束 raise StopIteration except Exception as e: # 处理其他异常,可根据需求调整 print(f"处理数据出错: {e}") raise StopIteration def __iter__(self): return self def __next__(self): return self._next() # 测试示例 r = Sub("cat sample_jsonlines.txt") for dict_line in r: print(dict_line)
关键修改说明
- 切断递归循环:
set_gen()直接从self.p.stdout读取行,不再迭代自身,彻底避免递归调用。 - 简化文本处理:初始化
subprocess.Popen时添加text=True,直接读取字符串行,省去字节解码步骤。 - 精准异常处理:明确捕获
StopIteration处理生成器耗尽的情况,替换原代码中模糊的全局异常捕获,同时增加JSON解析错误的针对性处理。 - 数据清理:自动去除行尾换行符和空行,避免拼接JSON数组时出现语法错误。
内容的提问来源于stack exchange,提问作者user1179317
相关产品推荐
相关产品推荐

