编写pandas文件转换类时如何避免重复的逻辑代码?
消除重复代码的可行方案
方案1:抽取公共校验逻辑为私有方法(最常用、易维护)
把重复的类型判断逻辑封装到内部私有方法中,不同转换逻辑的差异项(转换函数、存储路径、专属参数)作为入参传入即可。
代码示例:
from dataclasses import dataclass from typing import Union, Callable, Any import pandas as pd @dataclass class Converter: data: Union[str, pd.DataFrame] def _convert(self, convert_func: Callable, file_path: str, **kwargs): # 公共校验逻辑仅实现一次 if isinstance(self.data, pd.DataFrame): return convert_func(self.data, file_path, **kwargs) return self.data def to_pickle(self): """Check the type of data. Returns: a pickle file if the table exists, a string otherwise. """ return self._convert(pd.to_pickle, "table_data.pkl") def to_csv(self): """Check the type of data. Returns: a csv file if the table exists, a string otherwise. """ return self._convert( pd.to_csv, "./table_data.csv", index=False, encoding="utf_8_sig" )
优点:逻辑直观,后续新增转换方法只需要调用_convert传入差异参数即可,维护成本极低。
方案2:使用装饰器封装校验逻辑
如果后续要扩展的转换方法很多,还可以用装饰器把校验逻辑抽离出来,无需修改类的核心结构:
代码示例:
from functools import wraps from dataclasses import dataclass from typing import Union import pandas as pd def df_check(func): @wraps(func) def wrapper(self, *args, **kwargs): if isinstance(self.data, pd.DataFrame): return func(self, *args, **kwargs) return self.data return wrapper @dataclass class Converter: data: Union[str, pd.DataFrame] @df_check def to_pickle(self): """Check the type of data. Returns: a pickle file if the table exists, a string otherwise. """ return pd.to_pickle(self.data, "table_data.pkl") @df_check def to_csv(self): """Check the type of data. Returns: a csv file if the table exists, a string otherwise. """ return pd.to_csv( self.data, "./table_data.csv", index=False, encoding="utf_8_sig",)
优点:校验逻辑和业务转换逻辑完全解耦,装饰器可以复用在其他类的同类校验场景中。
方案3:配置化动态生成转换方法
如果转换类型非常多,还可以提前把所有转换规则做成配置字典,批量生成类方法,适合需要支持数十种格式转换的场景:
from dataclasses import dataclass from typing import Union import pandas as pd @dataclass class Converter: data: Union[str, pd.DataFrame] # 提前定义所有转换规则 convert_rules = { "to_pickle": { "func": pd.to_pickle, "path": "table_data.pkl", "kwargs": {} }, "to_csv": { "func": pd.to_csv, "path": "./table_data.csv", "kwargs": {"index": False, "encoding": "utf_8_sig"} } } # 动态批量生成方法 for method_name, rule in convert_rules.items(): def generate_method(rule): def _method(self): if isinstance(self.data, pd.DataFrame): return rule["func"](self.data, rule["path"], **rule["kwargs"]) return self.data _method.__name__ = method_name # 此处可手动给方法赋值docstring,和原有注释保持一致 return _method setattr(Converter, method_name, generate_method(rule))
优点:新增转换格式只需要加配置项,不需要重复写方法定义代码。
内容的提问来源于stack exchange,提问作者retr0327
相关产品推荐
相关产品推荐

