生成器读取大文件时,readline()跳过表头会导致每次next漏数据吗?
生成器调用
next()时是否会遗漏数据? 原代码示例
def read_large_file(file_object): # Define read_large_file() """A generator function to read a large file lazily.""" file_object.readline() # Skip the header line # Loop indefinitely until the end of the file while True: # Read a line from the file: data data = file_object.readline() # Break if this is the end of the file if not data: break # Yield the line of data yield data with open('world_dev_ind.csv') as file: #Open a connection to the file # Create a generator object for the file: gen_file gen_file = read_large_file(file) # Print the first three lines of the file print(next(gen_file)) print(next(gen_file)) print(next(gen_file))
问题
在上述代码中,如下语句:
# Skip the header line file_object.readline()每次对生成器对象调用
next()时,是否会遗漏一行数据?
解答
不会每次调用next()都遗漏数据,具体原因:
- 生成器函数
read_large_file在首次创建生成器对象时,会执行file_object.readline()这行代码,一次性跳过文件的表头行,这段代码只会执行一次,不会在每次next()调用时重复运行。 - 后续每次调用
next(gen_file),程序会进入while True循环逻辑:读取一行数据,判断是否到文件末尾,若未到则返回该行数据。整个过程不会再触发表头行的跳过操作。 - 实际运行时,假设文件表头是第1行,第一次
next()返回第2行,第二次返回第3行,第三次返回第4行,完全符合预期,没有重复遗漏数据的情况。
内容的提问来源于stack exchange,提问作者Sakshi
相关产品推荐
相关产品推荐

