You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何匹配文件内特定模式并切割内容生成独立文件?

解决方案

你可以通过状态机遍历文件内容,精准识别每个AV块的起始和结束位置,再将每个块写入对应的文件。以下提供两种实现方式:

Python 实现(适合脚本场景)

这个方法用状态机逐步匹配块的结构:先找到连续的「全#行 → 带AV的#行 → 全#行」作为块开头,然后收集内容直到下一个全#行,完成一个块的提取。

def is_all_hash(line):
    stripped = line.strip()
    return stripped and all(c == '#' for c in stripped)

def extract_av_blocks():
    # 读取输入文件所有行
    with open('input.txt', 'r', encoding='utf-8') as f:
        lines = [line.rstrip('\n') for line in f]
    
    blocks = []
    current_block = []
    state = 0  # 0: 寻找第一个全#行;1: 寻找带AV的行;2: 寻找第二个全#行;3: 收集内容直到结束全#行

    for line in lines:
        if state == 0:
            if is_all_hash(line):
                current_block.append(line)
                state = 1
        elif state == 1:
            if 'AV' in line and line.strip().startswith('#'):
                current_block.append(line)
                state = 2
            else:
                # 不符合AV行,重置状态
                current_block = []
                state = 0
        elif state == 2:
            if is_all_hash(line):
                current_block.append(line)
                state = 3
            else:
                current_block = []
                state = 0
        elif state == 3:
            current_block.append(line)
            if is_all_hash(line):
                # 块结束,加入列表并重置
                blocks.append(current_block)
                current_block = []
                state = 0

    # 将每个块写入对应文件
    for idx, block in enumerate(blocks, 1):
        with open(f'{idx}.txt', 'w', encoding='utf-8') as f:
            f.write('\n'.join(block) + '\n')

if __name__ == '__main__':
    extract_av_blocks()

Awk 实现(适合命令行快速处理)

如果你习惯用命令行工具,Awk可以直接处理文本,无需编写复杂脚本:

BEGIN { block_idx = 0; in_block = 0; hash_count = 0 }
/^#+$/ {
    if (!in_block) {
        hash_count++
        if (hash_count == 1) {
            first_hash = $0
        } else if (hash_count == 2) {
            # 检查上一行是否是AV标记行
            if (prev_line ~ /^#+ AV/) {
                block_idx++
                print first_hash > block_idx ".txt"
                print prev_line > block_idx ".txt"
                print $0 > block_idx ".txt"
                in_block = 1
            }
            hash_count = 0
        }
    } else {
        # 块内遇到结束的全#行,写入并退出块
        print $0 > block_idx ".txt"
        in_block = 0
    }
    next
}
{
    if (in_block) {
        print $0 > block_idx ".txt"
    }
    prev_line = $0
}

使用时直接在命令行运行:

awk -f extract_blocks.awk input.txt

逻辑说明

两种方法的核心都是:

  1. 识别块的起始标记:连续三行「全#行 → 含AV的#行 → 全#行」
  2. 从起始标记开始,收集所有行直到遇到下一个全#行,形成完整块
  3. 将每个块依次写入1.txt、2.txt、3.txt

内容的提问来源于stack exchange,提问作者user3285014

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 21:45:10