如何用Python提取文件中含test的文本并按行写入新文件?
问题:提取含指定关键词的行并保持换行格式
我有一个名为file1的文本文件,内容如下:
HelloWorldTestClass MyTestClass2 MyTestClass4 MyHelloWorld ApexClass * ApexTrigger Book__c CustomObject 56.0
需要将其中包含“test”(不区分大小写)的行输出到file2,预期输出:
HelloWorldTestClass MyTestClass2 MyTestClass4
我编写了如下Python代码:
import re import os file_contents1 = f'{os.getcwd()}/build/testlist.txt' file2_path = f'{os.getcwd()}/build/optestlist.txt' with open(file_contents1, 'r') as file1: file1_contents = file1.read() # print(file1_contents) # output = [file1_contents.strip() for line in file1_contents if "TestClass" in line] # # Use a regular expression pattern to match strings that contain "test" test_strings = [x for x in file1_contents.split("\n") if re.search(r"test", x, re.IGNORECASE)] # x = test_strings.strip("['t]") # # Print the result with open(file2_path, 'w') as file2: # write the contents of the first file to the second file for test in test_strings: file2.write(test)
但运行后输出内容挤在一起:
HelloWorldTestClass MyTestClass2 MyTestClass4
不符合预期的按行格式,该如何解决?
解决方案
问题原因
使用file1.read()读取整个文件后,用split("\n")拆分的每个行元素不包含换行符,直接写入文件时就不会自动换行,导致所有内容挤在一行。
修复方案1:写入时添加换行符
在循环写入每个匹配行时,额外写入换行符"\n":
import re import os file_contents1 = f'{os.getcwd()}/build/testlist.txt' file2_path = f'{os.getcwd()}/build/optestlist.txt' with open(file_contents1, 'r') as file1: file1_contents = file1.read() test_strings = [x for x in file1_contents.split("\n") if re.search(r"test", x, re.IGNORECASE)] with open(file2_path, 'w') as file2: for test in test_strings: # 写入行内容+换行符 file2.write(test + "\n")
修复方案2:按行读取文件(更高效)
直接逐行读取文件,避免一次性读取整个内容,同时保留原始换行符:
import re import os file_contents1 = f'{os.getcwd()}/build/testlist.txt' file2_path = f'{os.getcwd()}/build/optestlist.txt' with open(file_contents1, 'r') as file1, open(file2_path, 'w') as file2: for line in file1: # 检查当前行是否包含test(不区分大小写) if re.search(r"test", line, re.IGNORECASE): file2.write(line)
这种方式不仅能保持原始换行格式,还能处理大文件时节省内存,代码逻辑也更简洁。
内容的提问来源于stack exchange,提问作者Neetesshhr
相关产品推荐
相关产品推荐

