Flutter开发如何提取epub文件文本内容对接text to speech功能
Flutter 提取EPUB纯文本实现方案
EPUB本质是标准ZIP压缩包,无需依赖重型EPUB解析库,按格式规范逐层解压解析即可提取纯文本,完全适配Text To Speech的输入要求。
依赖配置
在pubspec.yaml中添加两个轻量依赖,用于解压和XML解析:
dependencies: archive: ^3.4.10 xml: ^6.5.0
实现逻辑
- 解压EPUB压缩包:将EPUB文件读取为字节流后用ZIP解码器解压,拿到包内所有文件列表
- 定位内容索引:解析包内固定路径的
META-INF/container.xml文件,拿到OPF格式的内容描述文件路径;再解析OPF文件,按阅读顺序拿到所有章节对应的HTML/XHTML文件路径 - 清洗提取纯文本:读取所有章节文件内容,剔除标签、样式、脚本内容,合并多余空白字符,拼接为连续纯文本
核心实现代码
import 'dart:io'; import 'package:archive/archive.dart'; import 'package:xml/xml.dart'; Future<String> extractEpubPlainText(String epubPath) async { final epubBytes = await File(epubPath).readAsBytes(); final archive = ZipDecoder().decodeBytes(epubBytes); // 解析container.xml获取OPF文件路径 String? opfPath; for (final file in archive) { if (file.name.toLowerCase() == 'meta-inf/container.xml') { final containerXml = XmlDocument.parse(String.fromCharCodes(file.content as List<int>)); opfPath = containerXml.findAllElements('rootfile').first.getAttribute('full-path'); break; } } if (opfPath == null) throw Exception('EPUB文件格式损坏'); // 解析OPF文件获取所有章节路径 List<String> chapterPaths = []; final opfDir = opfPath.contains('/') ? opfPath.substring(0, opfPath.lastIndexOf('/')) : ''; for (final file in archive) { if (file.name == opfPath) { final opfXml = XmlDocument.parse(String.fromCharCodes(file.content as List<int>)); final manifest = <String, String>{}; opfXml.findAllElements('item').forEach((item) { final id = item.getAttribute('id'); final href = item.getAttribute('href'); if (id != null && href != null) manifest[id] = href; }); opfXml.findAllElements('itemref').forEach((itemref) { final idref = itemref.getAttribute('idref'); if (idref != null && manifest.containsKey(idref)) { final relativePath = manifest[idref]!; final fullPath = relativePath.startsWith('/') ? relativePath.substring(1) : '$opfDir/$relativePath'.replaceAll('./', ''); chapterPaths.add(fullPath); } }); break; } } // 提取并清洗章节文本 final fullText = StringBuffer(); final tagRegex = RegExp(r'<[^>]*>', multiLine: true, caseSensitive: false); final whitespaceRegex = RegExp(r'\s+', multiLine: true); final styleScriptRegex = RegExp(r'<(style|script)[\s\S]*?</\1>', caseSensitive: false); for (final file in archive) { if (chapterPaths.contains(file.name) && file.isFile) { final rawContent = String.fromCharCodes(file.content as List<int>); final cleanText = rawContent .replaceAll(styleScriptRegex, '') .replaceAll(tagRegex, '') .replaceAll(whitespaceRegex, ' ') .trim(); if (cleanText.isNotEmpty) fullText.writeln(cleanText); } } return fullText.toString().trim(); }
TTS适配优化
- 可按章节拆分文本,支持TTS分段朗读、进度跳转,避免一次性加载大文件占用过多内存
- 可通过关键词过滤版权页、目录、脚注等非正文内容,减少无效朗读内容
- 带DRM版权加密的EPUB文件无法通过该方法解压,这类文件受版权保护无通用解密方案
内容的提问来源于stack exchange,提问作者Harsh Vardhan
相关产品推荐
相关产品推荐

