如何从大文本文件指定行读取,及崩溃后从记录行号重启
Got it, let's solve this problem of efficiently resuming large text file parsing without re-reading from the start every time—super frustrating when crashes make you lose progress, right?
核心思路:用字节偏移量替代纯行号定位
The big issue with just tracking line numbers is that text files don't have built-in line indexes. To jump to line N, you'd have to read every line before it, which is agonizingly slow for huge files. Instead, we'll track byte offsets (the exact position in the file where the next line starts) alongside line numbers—this lets us jump straight to the last processed spot in O(1) time.
完整实现代码(结合你的现有片段)
Here's how to modify your code to add crash-resilient, fast resuming:
import os # 配置项 PROGRESS_FILE = "parse_progress.txt" LARGE_TEXT_FILE = "your_target_file.txt" # 替换成你的超大文件路径 PARSED_DIR = "urltextdir/" SAVE_INTERVAL = 100 # 每处理100行保存一次进度,可根据需求调整 # 初始化进度 start_offset = 0 start_line = 0 all_parsed = os.listdir(PARSED_DIR) # 读取上次保存的进度 if os.path.exists(PROGRESS_FILE): with open(PROGRESS_FILE, "r") as pf: progress_data = pf.readline().strip() if progress_data: start_offset, start_line = map(int, progress_data.split(",")) count = start_line # 从记录的行号开始计数 with open(LARGE_TEXT_FILE, "r") as f: # 直接跳转到上次中断的字节位置,无需从头遍历 f.seek(start_offset) for line in f: count += 1 # 这里放你的逐行操作逻辑,比如检查是否已解析 # 示例:if line.strip() not in all_parsed: # # 执行你的处理代码... # 定期保存进度,减少IO开销 if count % SAVE_INTERVAL == 0: current_offset = f.tell() # 获取当前文件指针的字节偏移量 with open(PROGRESS_FILE, "w") as pf: pf.write(f"{current_offset},{count}") # 处理完成后清理进度文件(可选) if os.path.exists(PROGRESS_FILE): os.remove(PROGRESS_FILE)
关键细节说明
f.seek(start_offset): This is the magic line—instead of reading every line from the start, we jump directly to the exact byte position where we left off. Instant positioning, no wasted time.f.tell(): Gets the current byte position of the file pointer right after processing a line. We save this so we know where to resume next time.- Progress saving interval: Saving after every line would slow things down with constant disk writes. Saving every 100 lines strikes a balance—if a crash happens, you only lose up to 100 lines of work.
- Dual progress tracking: We save both the byte offset (for positioning) and line number (for your progress counting) to cover all bases.
注意事项
- If your large text file gets modified (added/removed lines) between sessions, this method will break—since the byte offsets will no longer correspond to the correct lines. Make sure the file is static during parsing.
- For multi-byte encodings like UTF-8,
seekworks fine here becausef.tell()returns positions at line boundaries (after the newline character), so we never land in the middle of a character. - If you absolutely need to jump to a specific line number (not just resume), you can pre-generate a line index file on first run—mapping each line number to its byte offset. This way, future jumps are instant, but you'll have to do one full pass upfront to build the index.
内容的提问来源于stack exchange,提问作者Kenstars
相关产品推荐
相关产品推荐

