Python逐行更新文件问题:修改内容被追加至末尾而非替换行
问题:逐行更新文件特定行内容时,修改内容被追加到文件末尾而非替换
作为作业任务的一部分,要求在不将整个文件加载到内存的情况下,仅逐行读取来更新特定行。当前代码存在问题:修改后的内容会被写入文件末尾,而非替换对应的目标行。尝试使用seek方法但持续报错,当前代码是最接近可行的版本。期望实现的效果为:调用
heap.update('currency', 'PKR', '123')时,保留所有行的原有顺序,仅将所有行中的PKR替换为123。
现有代码
update方法
def update(self, col_name, old_value, new_value, ): if self.is_empty() == 0: #check if file is empty return i = self.findIndex(col_name) #find what column im comparing to if i == -1: #if column out of bound return with open(self.name, "r") as f: #open file for read mode while True: temp='' line = f.readline() #read line if line == '': #if line is empty, we are done break #split line into the 4 columns keep1 = line.split(',', 3)[0]; keep2 = line.split(',', 3)[1]; keep3 = line.split(',', 3)[2]; keep4 = line.split(',', 3)[3]; x = line.split(',')[i]; if (x == old_value): #compare value and check if its equle, if it is: #create a new string with value if i == 0: temp = new_value + "," + keep2 + "," + keep3 + "," + keep4; elif i == 1: temp = keep1 + "," + new_value + "," + keep3 + "," + keep4; elif i == 2: temp = keep1 + "," + keep2 + "," + new_value + "," + keep4; elif i == 3: temp = keep1 + "," + keep2 + "," + keep3 + "," + new_value; with open(self.name, "a") as ww: #write value into file ww.write(temp)
__init__方法
def __init__(self, file_name): self.name = file_name; f = open(self.name, "w"); f.close();
findIndex方法
def findIndex(self, col_name): if self.is_empty() == 0: return -1 elif col_name == 'lid': return 0 elif col_name == 'loan_amount': return 1 elif col_name == "currency": return 2 elif col_name == "sector": return 3
示例文本
lid,loan_amount,currency,sector 653051,300.0,PKR,Food 653053,575.0,PKR,Trns 653068,150.0,INR,Trns 653063,200.0,PKR,Arts 653084,400.0,PKR,Food 653067,200.0,INR,Agri 653078,400.0,PKR,Serv 653082,475.0,PKR,Manu 653048,625.0,PKR,Food 653060,200.0,PKR,Trns 653088,400.0,PKR,Sale 653089,400.0,PKR,Reta 653062,400.0,PKR,Clth 653075,225.0,INR,Agri 653054,300.0,PKR,Trns 653091,400.0,PKR,Reta 653052,875.0,PKR,Serv 653066,250.0,INR,Serv 653080,475.0,PKR,Serv 653065,250.0,PKR,Food 653055,350.0,PKR,Food 653050,575.0,PKR,Clth 653079,350.0,PKR,Arts 653061,250.0,PKR,Food 653074,250.0,INR,Agri 653069,250.0,INR,Cons 653056,475.0,PKR,Trns 653071,125.0,INR,Agri 653073,250.0,INR,Agri 653059,250.0,PKR,Clth 653087,400.0,PKR,Manu 653076,450.0,PKR,Reta
问题出在哪?
咱们一条条拆解:
- 写入模式完全错误:你用
"a"(追加模式)打开文件,这就相当于每次都在文件末尾追加内容,根本不会替换原来的行。而且你还在循环里重复打开文件追加,不仅效率极低,还会导致修改过的行被多次写入。 - 遗漏了不需要修改的行:只有找到匹配
old_value的行时才生成temp写入,其他行完全没处理,结果就是原文件内容保留,修改后的行又被加在后面,整个文件彻底混乱。 - 对文本文件修改的认知误区:文本文件没法直接原地修改特定行——因为每行的字节长度可能变化(比如这次
PKR改123长度一致,但如果改成长度不同的内容,后面的字节会直接错位)。正确的做法是读一行处理一行,写到临时文件,最后用临时文件替换原文件。 - 代码冗余且有潜在bug:你每次拆分行都重复调用
split(',',3)四次,完全没必要,拆一次存起来复用即可;另外line.split(',')[i]可能因为行尾的换行符导致匹配失败(比如最后一列会带换行),最好先把换行符去掉再处理。
修正后的实现
下面是符合需求的代码,既不加载整个文件到内存,又能正确替换目标行:
import os def update(self, col_name, old_value, new_value): if self.is_empty(): # 假设is_empty()返回True表示文件为空,可根据你的实际实现调整 return i = self.findIndex(col_name) if i == -1: return # 创建临时文件保存处理后的内容,避免直接修改原文件导致数据丢失 temp_file_name = self.name + ".tmp" # 同时打开原文件(读)和临时文件(写) with open(self.name, "r") as f_in, open(temp_file_name, "w") as f_out: for line in f_in: line = line.rstrip('\n') # 去掉行尾换行符,避免干扰匹配 if not line: f_out.write('\n') continue # 只拆分一次列,提升效率 cols = line.split(',', 3) # 防止列数不足导致索引错误 if len(cols) <= i: f_out.write(line + '\n') continue # 匹配到目标值则替换,否则直接保留原行 if cols[i] == old_value: cols[i] = new_value new_line = ','.join(cols) f_out.write(new_line + '\n') else: f_out.write(line + '\n') # 用临时文件替换原文件,完成更新 os.replace(temp_file_name, self.name)
额外优化建议
你的__init__方法用"w"模式打开文件会直接清空已有内容,建议调整为仅在文件不存在时创建空文件:
def __init__(self, file_name): self.name = file_name # 只在文件不存在时创建空文件,避免覆盖已有数据 if not os.path.exists(file_name): with open(file_name, "w") as f: pass
内容的提问来源于stack exchange,提问作者Itzik Dan
相关产品推荐
相关产品推荐

