You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在指定起始索引后查找目标发票结束索引?

问题与解决方案

问题描述

需要返回目标发票起始索引之后的下一个invoice_index_end,但当前代码会错误找到起始索引之前的发票结束标记。希望通过循环遍历,定位到大于起始索引的目标结束索引。

原代码

lst_file = []
def assign_file_lines_to_list():
  for invoices in file:
    lst_file.append(invoices)
  return lst_file
assign_file_lines_to_list()
file.close()

#Opens the requested error log
error_lst = []
file_name = input("Enter the full name of the error log. \n")
error_file = open(file_name, "r")

#Putserror log file into a list word by word.
for eachWord in error_file:
  error_lst.extend(eachWord.split())

#Converts all string values to integers
def str_to_int():
  for i in range(0, len(error_lst)):
    try:
      error_lst[i] = int(error_lst[i])
    except:
      continue
  return error_lst

#Everything that is converted to an integer is added to a new list
int_lst = []
for eachword in str_to_int():
  if type(eachword) == int and eachword > 999:
    int_lst.append(eachword)

#And then turned back into a string.
def int_to_str():
  for i in range(0, len(int_lst)):
    try:
      int_lst[i] = str(int_lst[i])
    except:
      print("Error converting integer to a string!")
int_to_str()

#Standardizes all invoice to 10 digits
str_lst = [str(item).zfill(10) for item in int_lst]
print(str_lst)

#Finds the index of the invoice start and invoice end
new_lst = []
invoice_index_start = 0
invoice_index_end = 0
constant = '</Invoice>\n'
for i in str_lst: #integer
  invoice_index_start = lst_file.index('<InvoiceNumber>' + i + '</InvoiceNumber>\n')
  while str_lst.index(constant) > invoice_index_start:
    invoice_index_end = lst_file.index(constant)
  #if invoice_index_end <= invoice_index_start:
#Copies everything between start and end index into new list to be deleted from original list later
    new_lst += lst_file[(invoice_index_start - 1):(invoice_index_end + 1)]

解决方案

核心问题是lst_file.index()仅返回第一个匹配项的索引,导致起始索引前的结束标记被错误匹配。解决思路是从起始索引的下一个位置开始,向后遍历查找第一个目标结束标记:

修改后的关键代码段

#Finds the index of the invoice start and invoice end
new_lst = []
# 根据文件实际内容调整常量,若原文件是转义后的标签则用'&lt;/Invoice&gt;\n'
constant = '</Invoice>\n'
for invoice_num in str_lst:
    # 构建当前发票的起始行标识
    start_line = '<InvoiceNumber>' + invoice_num + '</InvoiceNumber>\n'
    try:
        invoice_index_start = lst_file.index(start_line)
    except ValueError:
        print(f"发票号 {invoice_num} 未找到,跳过")
        continue
    
    # 从起始索引后一位开始,查找第一个结束标记
    invoice_index_end = -1
    for idx in range(invoice_index_start + 1, len(lst_file)):
        if lst_file[idx] == constant:
            invoice_index_end = idx
            break
    
    if invoice_index_end == -1:
        print(f"发票号 {invoice_num} 未找到对应的结束标记,跳过")
        continue
    
    # 复制起始到结束区间的内容(包含前后行)
    new_lst += lst_file[invoice_index_start - 1 : invoice_index_end + 1]

说明

  1. 替换原有的index()直接调用,改为从invoice_index_start+1开始遍历,确保找到的是当前发票之后的第一个结束标记。
  2. 增加异常处理,避免因发票号不存在或无对应结束标记导致程序崩溃。
  3. 调整常量constant的格式,需与文件中实际的结束标签格式一致(原代码的&lt;是HTML转义字符,若文件中是原始XML标签则用<)。

内容的提问来源于stack exchange,提问作者phaynes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 03:20:14