Python读取HTML文件后列表迭代搜索的索引错误排查与解决
修正Python迭代搜索HTML文件时的索引计算错误
场景与问题
现有一个HTML文件,内容如下:
<span style="color: #000000; font-weight: bold; ">data 27 dic 2022 - 14:25:59::</span><br> <br> <span style="color: #0000ff; ">something</span> -> ALTO<br> <br> <span style="color: #000000; font-weight: bold; ">data 27 dic 2022 - 14:38:38::</span><br> <br> <span style="color: #0000ff; ">something</span> -> NO_ACQ<br> <br> <span style="color: #000000; font-weight: bold; ">data 27 dic 2022 - 14:38:40::</span><br>
需要通过Python的file.readlines()将其读取为列表后,迭代搜索包含data的行提取时间戳,且每次搜索无需从列表开头开始。当前实现的主代码和timestamp_search函数存在索引传递逻辑错误,导致迭代时索引混乱:
主代码
exit = 0 iterazione = 0 # 初始搜索起始索引 passed_time = ... # 业务定义的时间下限 actual_time = ... # 业务定义的时间上限 while exit == 0: timestamp, stop, index = timestamp_search(path, "data", iterazione) print('timestamp=', timestamp) if timestamp > passed_time and timestamp < actual_time: # do something(执行目标业务逻辑) exit = 1 elif stop == 1: exit = 1 else: iterazione = index + 1
原timestamp_search函数
def timestamp_search(file_path, word_1, iterazione): stop = 0 with open(file_path, 'r') as file: lst = file.readlines() print('iterazione', iterazione) if lst: if iterazione < len(lst): for index, line in enumerate(lst[iterazione:]): if line.find(word_1) != -1: print(index, line) # do something to calculate timestamp(时间戳提取逻辑) break else: stop = 1 return timestamp, stop else: print('File vuoto') return timestamp, stop, index
错误原因
使用enumerate(lst[iterazione:])时,返回的index是切片后的相对索引,而非原列表的绝对索引。例如:
- 第一次
iterazione=0,找到相对索引0,对应原列表索引0; - 第二次
iterazione=1,搜索lst[1:]时找到相对索引3,对应原列表索引应为1+3=4,但原代码直接返回相对索引3,导致下一次iterazione=3+1=4,后续搜索逻辑混乱。
修正方案
1. 计算绝对索引
在找到目标行的相对索引后,加上起始偏移量iterazione,得到原列表的绝对索引。
2. 完善变量初始化与边界处理
- 初始化
timestamp变量,避免未定义报错; - 处理遍历完切片后未找到目标行的情况;
- 修正返回值的完整性。
修正后的timestamp_search函数
def timestamp_search(file_path, word_1, iterazione): stop = 0 timestamp = None # 初始化timestamp变量 absolute_index = iterazione # 默认起始索引 with open(file_path, 'r') as file: lst = file.readlines() print('iterazione', iterazione) if lst: if iterazione < len(lst): # 遍历切片,同时计算原列表的绝对索引 for rel_index, line in enumerate(lst[iterazione:]): absolute_index = iterazione + rel_index if line.find(word_1) != -1: print(absolute_index, line) # 补充时间戳提取逻辑示例: time_str = line.split('>')[1].split('::')[0].strip() # 根据业务需求将time_str转换为timestamp格式,例如: # from datetime import datetime # timestamp = datetime.strptime(time_str, "data %d %b %Y - %H:%M:%S") break else: # 遍历完切片未找到目标行,标记终止 stop = 1 else: stop = 1 else: print('File vuoto') stop = 1 return timestamp, stop, absolute_index
修正后逻辑说明
- 每次搜索时,通过
iterazione + rel_index计算原列表的绝对索引,确保下一次搜索的起始位置正确; - 处理了遍历完切片未找到目标行的情况,避免索引异常;
- 初始化
timestamp变量,解决未定义的报错问题。
内容的提问来源于stack exchange,提问作者martinmistere
相关产品推荐
相关产品推荐

