如何使用Python删除表格文本文件中含空值或负数的行
问题描述
我正在处理.txt格式的表格文本文件,想要知道删除文件中包含负数值或空条目的行的最优方式是什么?
目前写的代码运行后仍然会写入所有文本文件条目,达不到仅保留无负数、无空条目有效行的要求。
原有问题代码
import os, sys inFile = sys.argv[1] baseN = os.path.basename(inFile) outFile = 'c:/exampleSolution.txt' #if path exists, read and write file if os.path.exists(inFile): inf = open(inFile,'r') outf = open(outFile,'w') #reading and writing header header = inf.readline() outf.write(header) not_consider = [] lines = inf.read().splitlines() for i in range(0,len(lines)): data = lines[i].split(' ') for j in range(0,len(data)): if (data[j] == '' or int(data[j]) < 0): #if line is having blank or negtive value # append i value to the not_consider list leaveOut.append(i) for i in range(0,len(lines)): #if i is in not_consider list, don't write to out file if i not in leaveOut: outf.write(lines[i]) print(lines[i]) outf.write("\n") inf.close() outf.close()
输入文件示例逻辑
样例表格中站点编号2和4存在负数/缺失项,需要删除,仅保留其他有效行。
排查结果与优化方案
原代码核心问题
- 变量名不统一:提前定义的待排除行列表名为
not_consider,实际添加排除行号时调用的是未定义的leaveOut变量,导致排除逻辑完全失效,所有行都被判定为有效 - 行分割逻辑错误:使用
split(' ')按单个空格分割,如果行内有多个连续空格、首尾空格,会生成大量无意义的空字符串,误判空条目 - 资源管理不规范:直接打开文件未用上下文管理器,异常场景下可能出现文件句柄泄漏
- 逻辑冗余:不需要先收集所有待排除行号再遍历写入,可以逐行判断直接写入,减少内存占用,大文件场景下效率更高
修正后可用代码
import os import sys def is_valid_line(line): # 按任意空白符分割行,自动忽略首尾空格和连续空格 items = line.strip().split() for item in items: # 判断空值 if not item: return False # 判断是否为负整数 try: num = int(item) if num < 0: return False except ValueError: # 若允许非整数字段可调整此处逻辑,默认非整数判定为无效行 return False return True if __name__ == "__main__": in_file = sys.argv[1] out_file = 'c:/exampleSolution.txt' if not os.path.exists(in_file): print(f"输入文件{in_file}不存在") sys.exit(1) # 用with上下文管理器自动管理文件句柄,无需手动关闭 with open(in_file, 'r', encoding='utf-8') as inf, open(out_file, 'w', encoding='utf-8') as outf: # 写入表头 header = inf.readline() outf.write(header) # 逐行处理剩余内容 for line in inf: if is_valid_line(line): outf.write(line) print(line.strip())
内容的提问来源于stack exchange,提问作者Dean OKeefe
相关产品推荐
相关产品推荐

