Python处理文本文件时.tell()方法不可靠的问题及替代方案问询
解决Python中.tell()不可靠及while循环检测EOF的问题
问题根源:文本模式下.tell()的"逻辑位置"陷阱
你遇到的超大数值(比如18446744073709551630)其实是文本模式下tell()的特性导致的:当以文本模式("r")打开文件时,tell()返回的不是实际的字节偏移量,而是内部维护的逻辑位置标记,这个值和你用file_len_chars计算的字符数完全不匹配——因为文本模式会自动转换换行符(比如Windows的\r\n会被转成\n),加上空行、特殊字符的处理,很容易出现位置计算混乱,甚至溢出成无符号的超大整数。
方案1:修复.tell()的使用(不推荐,但满足你的需求)
如果一定要坚持用tell(),唯一的办法是以二进制模式打开文件,这样tell()会返回准确的字节偏移量,和预先计算的文件字节数一致。但需要额外处理编码转换:
import os import re from random import randint # 保留你原有的辅助函数 def file_len_lines(f_name): with open(f_name) as f: for i, l in enumerate(f): pass return i + 1 def trim(sut): return re.sub(' +', ' ', sut).strip() # 测试文件生成逻辑不变 with open("test.txt", "w") as f: word_list = ("Betty Eats Cakes And Uncle Sells Eggs "*20).split() word_list[3] = "" for word in word_list: print(word, file=f) file_to_read = 'test.txt' # 修改文件打开方式为二进制模式 with open(file_to_read, "rb") as f: count = 0 file_length = os.path.getsize(file_to_read) file_length_lines = file_len_lines(file_to_read) print(f"Lines in file = {file_length_lines}, Bytes in file = {file_length}") f.seek(0) while f.tell() < file_length: count += 1 # 读取字节并解码为字符串 text_line = f.readline().decode('utf-8') print(f"Line = {count}, ", end="") print(f"Tell = {f.tell()}, ", end="") print(f"Length {len(text_line)} ", end="") if text_line in ['', '\n']: print(count) elif trim(text_line).upper()[0] in "A E I O U".split(): print(text_line, end='') else: print(count)
这里用os.path.getsize()直接获取文件字节数,比你自己写的file_len_chars更准确高效。
方案2:更可靠的while循环EOF检测(强烈推荐)
其实完全不用依赖tell()和预先计算文件长度,Python的readline()方法在到达EOF时会返回空字符串'',我们可以利用这一点来构建while循环,既简单又可靠,还避免了两次遍历文件的开销:
import re from random import randint # 保留你原有的辅助函数 def file_len_lines(f_name): with open(f_name) as f: for i, l in enumerate(f): pass return i + 1 def trim(sut): return re.sub(' +', ' ', sut).strip() # 测试文件生成逻辑不变 with open("test.txt", "w") as f: word_list = ("Betty Eats Cakes And Uncle Sells Eggs "*20).split() word_list[3] = "" for word in word_list: print(word, file=f) file_to_read = 'test.txt' with open(file_to_read, "r") as f: count = 0 file_length_lines = file_len_lines(file_to_read) print(f"Lines in file = {file_length_lines}") while True: count += 1 text_line = f.readline() # 检测EOF:readline返回空字符串时退出循环 if not text_line: break print(f"Line = {count}, ", end="") print(f"Length {len(text_line)} ", end="") if text_line in ['', '\n']: print(count) elif trim(text_line).upper()[0] in "A E I O U".split(): print(text_line, end='') else: print(count)
这个方式完全符合你的while循环需求,不用修改核心代码架构,而且彻底避免了tell()带来的所有问题,运行效率也更高(不用先遍历一遍文件计算长度)。
内容的提问来源于stack exchange,提问作者Lance Skelly
相关产品推荐
相关产品推荐

