如何在终端中从文本文件提取并复制关键词后的内容
扫描文本文件提取「Day X」类关键词后的内容
Python实现方案
核心逻辑
读取文本文件,匹配所有Day + 数字格式的关键词,提取每个关键词之后的文本,直到遇到下一个同类关键词或文件结束。
可运行代码
import re def get_content_after_day(file_path): # 匹配"Day 数字"的正则规则 day_pattern = re.compile(r'Day \d+') sections = [] current_keyword = None current_content = [] with open(file_path, 'r', encoding='utf-8') as f: for line in f: line_stripped = line.strip() # 检查当前行是否包含目标关键词 match = day_pattern.search(line) if match: # 先保存上一个关键词的内容(如果有的话) if current_keyword is not None: sections.append({ 'keyword': current_keyword, 'content': '\n'.join(current_content) }) # 切换到新的关键词 current_keyword = match.group() current_content = [] # 提取当前行关键词后的内容 after_part = line[match.end():].strip() if after_part: current_content.append(after_part) elif current_keyword is not None and line_stripped: # 收集当前关键词下的内容行 current_content.append(line_stripped) # 保存最后一个关键词的内容 if current_keyword is not None: sections.append({ 'keyword': current_keyword, 'content': '\n'.join(current_content) }) # 输出结果示例 for sec in sections: print(f"【{sec['keyword']}】") print(sec['content']) print('-' * 30) # 替换为你的文本文件路径 get_content_after_day('your_text_file.txt')
代码说明
- 正则表达式
Day \d+能匹配所有类似Day 1、Day 17的格式 - 逐行处理文件,自动切换关键词段落,避免遗漏内容
- 最终结果以结构化的方式存储,方便后续使用(比如写入新文件)
命令行快速方案(Linux/macOS)
如果不需要复杂处理,用grep和sed组合可以快速提取指定关键词后的内容:
# 提取Day 17之后的内容,直到下一个Day关键词出现 grep -A 2000 "Day 17" your_file.txt | sed '/Day [0-9]/q' | grep -v "Day 17"
-A 2000:获取匹配行之后的2000行(可根据文件大小调整数值)sed '/Day [0-9]/q':遇到下一个Day + 数字的行就停止输出grep -v "Day 17":移除关键词所在的行
内容的提问来源于stack exchange,提问作者Amey Deshpande
相关产品推荐
相关产品推荐

