子类继承场景下,如何让EpubTrad生成ChapterTrad实例而非Chapter?
现有结构
我正在对Epub电子书进行建模。Epub本质上是一个包含.html文档(作为章节)的.zip文件,我的类结构也遵循这一特点。
Epub类负责打开zip文件、查找并提取章节名称,同时创建Chapter实例列表。Chapter类负责读取文件、用BeautifulSoup解析并提取段落。
该结构运行良好,能满足我的所有需求。
以下是这两个类的极简版本(仅用于展示结构,不可直接使用):
class Epub: def __init__(self, zipped_file): # load the book in memory and fill in several attributes # [more things to do...] self.chap_file_names = [str(p) for p in self.chap_file_paths] # build a list of chapters self.chapters: list[Chapter] = [ Chapter(self.input_zip.read(chap_file_name)) for chap_file_name in self.chap_file_names ] class Chapter: def __init__(self, chap_content): # parse the soup and get the body self.soup = BeautifulSoup(chap_content, features="html.parser") self.body = self.soup.body # find the paragraphs self.all_p_tag = self.body.find_all("p")
扩展类
现在我需要扩展这两个类以添加额外功能,将它们命名为EpubTrad和ChapterTrad。
如果直接定义class EpubTrad(Epub):,chap_file_names等属性能正常继承;ChapterTrad也应继承自Chapter,因为soup等属性仍需保留使用。
但我遇到了问题:实例化EpubTrad对象时,chapters属性是Chapter实例列表,而我需要它是ChapterTrad实例列表。
该如何解决这个问题?
可行方案:拆分__init__
我尝试的一种解决方法是延迟章节列表的创建:
from typing import Union, IO, Path class Epub: def __init__( self, zipped_file: Union[str, IO[bytes], Path], ) -> None: # load the book in memory and fill in several attributes self.chap_file_names = [str(p) for p in self.chap_file_paths] def parse_chapters(self): # build a list of chapters self.chapters: list[Chapter] = [ Chapter(self.input_zip.read(chap_file_name)) for chap_file_name in self.chap_file_names ] class EpubTrad(Epub): def __init__( self, zipped_file: Union[str, IO[bytes], Path], lang: str, ) -> None: # load the book in memory and fill in epub attributes super().__init__(zipped_file) # do more init related to trad things self.lang = lang def parse_chapters(self): # build a list of trad chapters self.chapters: list[ChapterTrad] = [ ChapterTrad(self.input_zip.read(chap_file_name)) for chap_file_name in self.chap_file_names ] def cool_trad_func(self): """Do fancy trad-related things."""
使用方式如下:
>>> ep = Epub("book.epub") >>> ep.parse_chapters() >>> ept = EpubTrad("book.epub", "en") >>> ept.parse_chapters()
当然,如果__init__中解析章节后还有其他操作,需要放到finish_init这类额外方法中。
此外,Chapter和ChapterTrad也需要类似拆分,因为Chapter实际包含Paragraph实例列表,而该类也需扩展为ParagraphTrad,不过当前问题已足够复杂,暂不展开。
简而言之
Epub类的构造函数会创建Chapter实例列表。如何通过继承Epub得到EpubTrad类,使其生成继承自Chapter的ChapterTrad实例列表?
感谢!
内容的提问来源于stack exchange,提问作者Pietro

