流式数据填充@dataclass的优化方案及起始异常处理问询
问题
我有一组重复流式传输的数据,需要先应用业务逻辑再填充到@dataclass中。
数据类结构如下:
class Foo: color: str engine: int petrol: int diesel: int
数据包含枚举类型的color,以及两组两位数(以单个数字形式流式传输),示例数据如下:
Red, 3, 1, 4, 5, Blue, 2, 2, 7, 5, Orange, 5, 2, 6, 8 etc...
规则说明:
- 遇到
color枚举值(Red/Blue/Orange/Green/Yellow)时,赋值给Foo.color; - 紧跟
color的两个单个数字要拼接为两位数,赋值给Foo.engine; - 若
color为Red、Blue或Yellow,后续两个单个数字拼接为两位数赋值给Foo.petrol; - 若
color为Orange或Green,后续两个单个数字拼接为两位数赋值给Foo.diesel; - 数据可能从非
color枚举值开始传输,此时需跳过前1-2个值,直到遇到color枚举值再开始处理。
我之前的实现代码可读性差且逻辑混乱,代码如下:
digits = 0 positional_counter = 0 # Set first attribute color if value in ["Red", "Blue", "Orange", "Green"]: self.foo.color = value positional_counter += 1 # Set second attribute - joining the digits if positional_counter == 1: digits = value * 10 positional_counter += 1 if positional_counter == 2: self.foo.engine = digits + value digits = 0 positional_counter += 1 # Whether the second set of digits is set against petrol or diesel depends on color attrib if self.foo.color in ["Red", "Blue"]: if positional_counter == 3: digits = value * 10 positional_counter += 1 if positional_counter == 4: self.foo.petrol = digits + value positional_counter = 0 if self.foo.color in ["Orange", "Green"]: if positional_counter == 3: digits = value * 10 positional_counter += 1 if positional_counter == 4: self.foo.diesel = digits + value positional_counter = 0
请问有没有更优的实现方式?
优化实现方案
可以采用状态机模式处理流式数据,把每个处理阶段定义为明确的状态,逻辑清晰且能轻松处理起始无效数据的情况。
1. 定义状态枚举
先把处理过程中的状态明确出来,替代零散的计数器:
from dataclasses import dataclass from enum import Enum class ProcessingState(Enum): WAITING_FOR_COLOR = 0 # 等待合法color值 EXPECTING_ENGINE_DIGIT_1 = 1 # 等待engine的第一位数字 EXPECTING_ENGINE_DIGIT_2 = 2 # 等待engine的第二位数字 EXPECTING_FUEL_DIGIT_1 = 3 # 等待燃油属性的第一位数字 EXPECTING_FUEL_DIGIT_2 = 4 # 等待燃油属性的第二位数字
2. 封装处理器类
把状态、临时数据和目标数据封装到处理器类中,逻辑更集中:
@dataclass class Foo: color: str = "" engine: int = 0 petrol: int = 0 diesel: int = 0 class FooDataProcessor: def __init__(self): self.state = ProcessingState.WAITING_FOR_COLOR self.current_foo = Foo() self.temp_digit = 0 # 临时存储拼接数字的第一位 def process_value(self, value): # 处理字符串类型的数字 if isinstance(value, str) and value.isdigit(): value = int(value) # 根据当前状态分发处理逻辑 if self.state == ProcessingState.WAITING_FOR_COLOR: self._handle_waiting_for_color(value) elif self.state == ProcessingState.EXPECTING_ENGINE_DIGIT_1: self._handle_engine_digit1(value) elif self.state == ProcessingState.EXPECTING_ENGINE_DIGIT_2: self._handle_engine_digit2(value) elif self.state == ProcessingState.EXPECTING_FUEL_DIGIT_1: self._handle_fuel_digit1(value) elif self.state == ProcessingState.EXPECTING_FUEL_DIGIT_2: self._handle_fuel_digit2(value) def _handle_waiting_for_color(self, value): # 只处理合法color,其他值直接跳过 valid_colors = {"Red", "Blue", "Orange", "Green", "Yellow"} if value in valid_colors: self.current_foo = Foo(color=value) # 重置新的Foo实例 self.state = ProcessingState.EXPECTING_ENGINE_DIGIT_1 def _handle_engine_digit1(self, value): if isinstance(value, int): self.temp_digit = value * 10 self.state = ProcessingState.EXPECTING_ENGINE_DIGIT_2 def _handle_engine_digit2(self, value): if isinstance(value, int): self.current_foo.engine = self.temp_digit + value self.state = ProcessingState.EXPECTING_FUEL_DIGIT_1 def _handle_fuel_digit1(self, value): if isinstance(value, int): self.temp_digit = value * 10 self.state = ProcessingState.EXPECTING_FUEL_DIGIT_2 def _handle_fuel_digit2(self, value): if isinstance(value, int): fuel_value = self.temp_digit + value # 根据color赋值对应燃油属性 if self.current_foo.color in {"Red", "Blue", "Yellow"}: self.current_foo.petrol = fuel_value elif self.current_foo.color in {"Orange", "Green"}: self.current_foo.diesel = fuel_value # 处理完一组数据后,可在此触发后续操作(如保存/使用实例) print(f"处理完成的Foo实例: {self.current_foo}") # 回到初始状态,准备处理下一组数据 self.state = ProcessingState.WAITING_FOR_COLOR
3. 使用示例
# 模拟流式数据输入 stream_data = ["X", 3, "Red", 3, 1, 4, 5, "Blue", 2, 2, 7, 5, "Orange", 5, 2, 6, 8] processor = FooDataProcessor() for val in stream_data: processor.process_value(val)
优化点说明
- 状态清晰:每个状态对应明确的处理逻辑,避免了零散的计数器判断,可读性和维护性大幅提升;
- 自动跳过无效数据:在
WAITING_FOR_COLOR状态下,所有非合法color值都会被直接跳过,无需额外处理; - 模块化:把每个状态的处理逻辑拆分为独立方法,代码结构更清晰;
- 扩展性强:如果后续业务规则变化(比如新增color类型、新增属性),只需新增状态或修改对应处理方法即可。
内容的提问来源于stack exchange,提问作者Al Grant
相关产品推荐
相关产品推荐

