You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

编写pandas文件转换类时如何避免重复的逻辑代码?

消除重复代码的可行方案

方案1:抽取公共校验逻辑为私有方法(最常用、易维护)

把重复的类型判断逻辑封装到内部私有方法中,不同转换逻辑的差异项(转换函数、存储路径、专属参数)作为入参传入即可。
代码示例:

from dataclasses import dataclass
from typing import Union, Callable, Any
import pandas as pd

@dataclass
class Converter:
    data: Union[str, pd.DataFrame]

    def _convert(self, convert_func: Callable, file_path: str, **kwargs):
        # 公共校验逻辑仅实现一次
        if isinstance(self.data, pd.DataFrame):
            return convert_func(self.data, file_path, **kwargs)
        return self.data

    def to_pickle(self):
        """Check the type of data. 

        Returns:
            a pickle file if the table exists, a string otherwise.
        """
        return self._convert(pd.to_pickle, "table_data.pkl")
        
    def to_csv(self):
        """Check the type of data. 

        Returns:
            a csv file if the table exists, a string otherwise.
        """
        return self._convert(
            pd.to_csv, "./table_data.csv", 
            index=False, 
            encoding="utf_8_sig"
        )

优点:逻辑直观,后续新增转换方法只需要调用_convert传入差异参数即可,维护成本极低。

方案2:使用装饰器封装校验逻辑

如果后续要扩展的转换方法很多,还可以用装饰器把校验逻辑抽离出来,无需修改类的核心结构:
代码示例:

from functools import wraps
from dataclasses import dataclass
from typing import Union
import pandas as pd

def df_check(func):
    @wraps(func)
    def wrapper(self, *args, **kwargs):
        if isinstance(self.data, pd.DataFrame):
            return func(self, *args, **kwargs)
        return self.data
    return wrapper

@dataclass
class Converter:
    data: Union[str, pd.DataFrame]

    @df_check
    def to_pickle(self):
        """Check the type of data. 

        Returns:
            a pickle file if the table exists, a string otherwise.
        """
        return pd.to_pickle(self.data, "table_data.pkl")
        
    @df_check
    def to_csv(self):
        """Check the type of data. 

        Returns:
            a csv file if the table exists, a string otherwise.
        """
        return pd.to_csv(
            self.data, "./table_data.csv", 
            index=False, 
            encoding="utf_8_sig",)

优点:校验逻辑和业务转换逻辑完全解耦,装饰器可以复用在其他类的同类校验场景中。

方案3:配置化动态生成转换方法

如果转换类型非常多,还可以提前把所有转换规则做成配置字典,批量生成类方法,适合需要支持数十种格式转换的场景:

from dataclasses import dataclass
from typing import Union
import pandas as pd

@dataclass
class Converter:
    data: Union[str, pd.DataFrame]

# 提前定义所有转换规则
convert_rules = {
    "to_pickle": {
        "func": pd.to_pickle,
        "path": "table_data.pkl",
        "kwargs": {}
    },
    "to_csv": {
        "func": pd.to_csv,
        "path": "./table_data.csv",
        "kwargs": {"index": False, "encoding": "utf_8_sig"}
    }
}

# 动态批量生成方法
for method_name, rule in convert_rules.items():
    def generate_method(rule):
        def _method(self):
            if isinstance(self.data, pd.DataFrame):
                return rule["func"](self.data, rule["path"], **rule["kwargs"])
            return self.data
        _method.__name__ = method_name
        # 此处可手动给方法赋值docstring,和原有注释保持一致
        return _method
    setattr(Converter, method_name, generate_method(rule))

优点:新增转换格式只需要加配置项,不需要重复写方法定义代码。

内容的提问来源于stack exchange,提问作者retr0327

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 19:45:03