如何用Python在Jupyter中查找文本指定词,匹配行追加写入新文件
Jupyter环境下文本关键词筛选追加方案
以下两种实现方式均支持大文件处理、多文件批量处理、追加写入不覆盖原有内容:
前置依赖(使用pandas方案需要先安装)
在Jupyter单元格中运行以下命令安装pandas:
!pip install pandas
方案1:pandas分块读取(适配指定技术栈)
核心是用pandas的chunksize参数分块读取大文件,避免一次性加载全量内容导致内存溢出,筛选后以追加模式写入目标文件:
import pandas as pd # -------------------------- 自定义配置项 -------------------------- search_keyword = "test" # 要查找的关键词 source_files = ["./demo1.txt", "./demo2.txt"] # 待检索的源文件路径,支持多个 target_file = "./matched_result.txt" # 结果输出的目标文件路径 chunk_size = 10000 # 每次读取的行数,可根据内存大小调整 # -------------------------- 执行逻辑 -------------------------- for file_path in source_files: # 分块读取源文件,默认按行读取,不设表头 for chunk in pd.read_csv(file_path, header=None, chunksize=chunk_size, dtype=str, sep="\n", keep_default_na=False): # 筛选包含目标关键词的行 matched_lines = chunk[chunk[0].str.contains(search_keyword, na=False)] # 追加写入目标文件,不写索引、不写表头 matched_lines.to_csv(target_file, mode="a", header=False, index=False, encoding="utf-8")
方案2:纯Python逐行读取(内存占用最低,适合超大体积文件)
不需要依赖第三方库,逐行读取逐行判断,内存占用几乎可以忽略:
# -------------------------- 自定义配置项 -------------------------- search_keyword = "test" source_files = ["./demo1.txt", "./demo2.txt"] target_file = "./matched_result.txt" # -------------------------- 执行逻辑 -------------------------- for file_path in source_files: with open(file_path, "r", encoding="utf-8") as f_in: with open(target_file, "a", encoding="utf-8") as f_out: for line in f_in: if search_keyword in line: f_out.write(line)
功能说明
- 两种方案均支持单/多文件批量处理,匹配到的行按源文件顺序追加到目标文件末尾,不会覆盖目标文件原有内容
- 针对示例场景,输入文件处理后,目标文件会精准保留两行包含
test的内容 - 如果需要忽略大小写匹配,可以把判断条件替换为
search_keyword.lower() in line.lower()(纯Python方案)或者添加case=False参数str.contains(search_keyword, case=False, na=False)(pandas方案)
内容的提问来源于stack exchange,提问作者scofx
相关产品推荐
相关产品推荐

