You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取小说正文,排除章节标题与卷首语(epigraphs)?

解决建议

方法一:Python 自定义脚本

这是适配性最强的方案,完全贴合你的章节结构:

  • 核心逻辑:逐行读取文本,先跳过匹配• CHAPTER [0-9]+ •格式的章节标题行,再跳过后续缩进的引用块(引号行和署名行),直到遇到非缩进的正文内容后开始保留。
  • 示例代码:
import re

def extract_clean_content(input_path, output_path):
    with open(input_path, 'r', encoding='utf-8') as f:
        lines = f.readlines()
    
    cleaned_lines = []
    skip_section = False
    skip_epigraph = False

    for line in lines:
        stripped = line.strip()
        # 匹配章节标题,开启跳过模式
        if re.match(r'^• CHAPTER \d+ •$', stripped):
            skip_section = True
            skip_epigraph = True
            continue
        # 跳过标题后的空行
        if skip_section and not stripped:
            continue
        # 跳过缩进的引用内容
        if skip_epigraph:
            if line.startswith(('    ', '         ')):
                continue
            else:
                skip_section = False
                skip_epigraph = False
                cleaned_lines.append(line)
        else:
            cleaned_lines.append(line)
    
    with open(output_path, 'w', encoding='utf-8') as f:
        f.writelines(cleaned_lines)

# 调用示例
extract_clean_content('raw_novel.txt', 'cleaned_novel.txt')
  • 用法:将原始小说文本保存为raw_novel.txt,运行脚本后会生成仅含正文的cleaned_novel.txt。

方法二:命令行工具(sed)

适合熟悉终端操作的用户,通过正则匹配批量过滤内容:

sed -E '/^• CHAPTER [0-9]+ •$/,/^[^ ]/{/^• CHAPTER [0-9]+ •$/d;/^    "/d;/^         -/d;/^$/d}' raw_novel.txt > cleaned_novel.txt
  • 逻辑说明:匹配从章节标题到第一个非缩进行的区间,删除标题行、缩进的引号行、署名行及空行,保留后续正文。

方法三:文本编辑器正则替换(VS Code/Subime Text)

无需编程,用编辑器自带功能批量清理:

  1. 打开小说文本,调出替换面板(VS Code按Ctrl+H,Sublime按Ctrl+Shift+H)
  2. 勾选「正则表达式」选项
  3. 查找内容输入:
^• CHAPTER \d+ •$\n(\n|    ".*"\n|         -.*\n)+
  1. 替换内容留空,点击「全部替换」
  • 逻辑:匹配章节标题及后续的空行、引用行、署名行,一次性替换为空,仅保留正文。

内容的提问来源于stack exchange,提问作者greenc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 10:51:10