You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

子类继承场景下,如何让EpubTrad生成ChapterTrad实例而非Chapter?

现有结构

我正在对Epub电子书进行建模。Epub本质上是一个包含.html文档(作为章节)的.zip文件,我的类结构也遵循这一特点。

Epub类负责打开zip文件、查找并提取章节名称,同时创建Chapter实例列表。Chapter类负责读取文件、用BeautifulSoup解析并提取段落。

该结构运行良好,能满足我的所有需求。

以下是这两个类的极简版本(仅用于展示结构,不可直接使用):

class Epub:
    def __init__(self, zipped_file):
        # load the book in memory and fill in several attributes
        # [more things to do...]
        self.chap_file_names = [str(p) for p in self.chap_file_paths]
        # build a list of chapters
        self.chapters: list[Chapter] = [
            Chapter(self.input_zip.read(chap_file_name))
            for chap_file_name in self.chap_file_names
        ]

class Chapter:
    def __init__(self, chap_content):
        # parse the soup and get the body
        self.soup = BeautifulSoup(chap_content, features="html.parser")
        self.body = self.soup.body
        # find the paragraphs
        self.all_p_tag = self.body.find_all("p")

扩展类

现在我需要扩展这两个类以添加额外功能,将它们命名为EpubTrad和ChapterTrad。

如果直接定义class EpubTrad(Epub):,chap_file_names等属性能正常继承;ChapterTrad也应继承自Chapter,因为soup等属性仍需保留使用。

但我遇到了问题:实例化EpubTrad对象时,chapters属性是Chapter实例列表,而我需要它是ChapterTrad实例列表。

该如何解决这个问题?

可行方案:拆分__init__

我尝试的一种解决方法是延迟章节列表的创建:

from typing import Union, IO, Path

class Epub:
    def __init__(
        self,
        zipped_file: Union[str, IO[bytes], Path],
    ) -> None:
        # load the book in memory and fill in several attributes
        self.chap_file_names = [str(p) for p in self.chap_file_paths]

    def parse_chapters(self):
        # build a list of chapters
        self.chapters: list[Chapter] = [
            Chapter(self.input_zip.read(chap_file_name))
            for chap_file_name in self.chap_file_names
        ]

class EpubTrad(Epub):
    def __init__(
        self,
        zipped_file: Union[str, IO[bytes], Path],
        lang: str,
    ) -> None:
        # load the book in memory and fill in epub attributes
        super().__init__(zipped_file)
        # do more init related to trad things
        self.lang = lang

    def parse_chapters(self):
        # build a list of trad chapters
        self.chapters: list[ChapterTrad] = [
            ChapterTrad(self.input_zip.read(chap_file_name))
            for chap_file_name in self.chap_file_names
        ]

    def cool_trad_func(self):
        """Do fancy trad-related things."""

使用方式如下:

>>> ep = Epub("book.epub")
>>> ep.parse_chapters()
>>> ept = EpubTrad("book.epub", "en")
>>> ept.parse_chapters()

当然,如果__init__中解析章节后还有其他操作,需要放到finish_init这类额外方法中。

此外,Chapter和ChapterTrad也需要类似拆分,因为Chapter实际包含Paragraph实例列表,而该类也需扩展为ParagraphTrad,不过当前问题已足够复杂,暂不展开。

简而言之

Epub类的构造函数会创建Chapter实例列表。如何通过继承Epub得到EpubTrad类,使其生成继承自Chapter的ChapterTrad实例列表?

感谢!


内容的提问来源于stack exchange,提问作者Pietro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 17:54:26