如何在指定起始索引后查找目标发票结束索引?
问题与解决方案
问题描述
需要返回目标发票起始索引之后的下一个invoice_index_end,但当前代码会错误找到起始索引之前的发票结束标记。希望通过循环遍历,定位到大于起始索引的目标结束索引。
原代码
lst_file = [] def assign_file_lines_to_list(): for invoices in file: lst_file.append(invoices) return lst_file assign_file_lines_to_list() file.close() #Opens the requested error log error_lst = [] file_name = input("Enter the full name of the error log. \n") error_file = open(file_name, "r") #Putserror log file into a list word by word. for eachWord in error_file: error_lst.extend(eachWord.split()) #Converts all string values to integers def str_to_int(): for i in range(0, len(error_lst)): try: error_lst[i] = int(error_lst[i]) except: continue return error_lst #Everything that is converted to an integer is added to a new list int_lst = [] for eachword in str_to_int(): if type(eachword) == int and eachword > 999: int_lst.append(eachword) #And then turned back into a string. def int_to_str(): for i in range(0, len(int_lst)): try: int_lst[i] = str(int_lst[i]) except: print("Error converting integer to a string!") int_to_str() #Standardizes all invoice to 10 digits str_lst = [str(item).zfill(10) for item in int_lst] print(str_lst) #Finds the index of the invoice start and invoice end new_lst = [] invoice_index_start = 0 invoice_index_end = 0 constant = '</Invoice>\n' for i in str_lst: #integer invoice_index_start = lst_file.index('<InvoiceNumber>' + i + '</InvoiceNumber>\n') while str_lst.index(constant) > invoice_index_start: invoice_index_end = lst_file.index(constant) #if invoice_index_end <= invoice_index_start: #Copies everything between start and end index into new list to be deleted from original list later new_lst += lst_file[(invoice_index_start - 1):(invoice_index_end + 1)]
解决方案
核心问题是lst_file.index()仅返回第一个匹配项的索引,导致起始索引前的结束标记被错误匹配。解决思路是从起始索引的下一个位置开始,向后遍历查找第一个目标结束标记:
修改后的关键代码段
#Finds the index of the invoice start and invoice end new_lst = [] # 根据文件实际内容调整常量,若原文件是转义后的标签则用'</Invoice>\n' constant = '</Invoice>\n' for invoice_num in str_lst: # 构建当前发票的起始行标识 start_line = '<InvoiceNumber>' + invoice_num + '</InvoiceNumber>\n' try: invoice_index_start = lst_file.index(start_line) except ValueError: print(f"发票号 {invoice_num} 未找到,跳过") continue # 从起始索引后一位开始,查找第一个结束标记 invoice_index_end = -1 for idx in range(invoice_index_start + 1, len(lst_file)): if lst_file[idx] == constant: invoice_index_end = idx break if invoice_index_end == -1: print(f"发票号 {invoice_num} 未找到对应的结束标记,跳过") continue # 复制起始到结束区间的内容(包含前后行) new_lst += lst_file[invoice_index_start - 1 : invoice_index_end + 1]
说明
- 替换原有的
index()直接调用,改为从invoice_index_start+1开始遍历,确保找到的是当前发票之后的第一个结束标记。 - 增加异常处理,避免因发票号不存在或无对应结束标记导致程序崩溃。
- 调整常量
constant的格式,需与文件中实际的结束标签格式一致(原代码的<是HTML转义字符,若文件中是原始XML标签则用<)。
内容的提问来源于stack exchange,提问作者phaynes
相关产品推荐
相关产品推荐

