使用while循环过滤列表重复项时遇索引越界错误的解决方法
问题
尝试打开txt文件并提取其中所有单词存入列表,再通过while循环过滤列表中的重复项。编写的代码运行时触发回溯错误,提示“list index is out of range(列表索引越界)”,将判断条件中的>1改为0后问题仍未解决,请问该如何修改代码以避免该错误?
代码:
fname = input("Enter file name: ") fh = open("romeo.txt") lst = list() for line in fh: nlst = (line.rstrip()).split() lst = lst + nlst for i in [*range(len(lst))]: if lst.count(lst[i]) > 1: while lst.count(lst[i]) > 1: print(lst[i]) lst.remove(lst[i]) else: continue print(lst)
错误原因
核心问题是for i in [*range(len(lst))]会在循环启动时生成一个基于原列表初始长度的索引序列,但循环过程中你通过lst.remove()不断缩短列表长度。当循环到后期,原索引值会超过当前列表的实际长度,直接触发索引越界错误。
解决方案
方案1:遍历列表副本,规避修改原列表的干扰
把遍历对象改成原列表的切片副本lst[:],这样遍历的是初始状态的元素,修改原列表不会影响遍历流程:
fname = input("Enter file name: ") fh = open("romeo.txt") lst = list() for line in fh: nlst = (line.rstrip()).split() lst = lst + nlst # 遍历列表副本,避免修改原列表导致的索引异常 for word in lst[:]: if lst.count(word) > 1: while lst.count(word) > 1: print(word) lst.remove(word) print(lst)
方案2:高效去重(保留元素首次出现顺序,Python3.7+)
如果不需要打印重复元素的过程,仅需得到去重后的列表,用Python3.7+自带的有序字典实现去重,效率远高于反复调用count()和remove():
fname = input("Enter file name: ") fh = open("romeo.txt") lst = list() for line in fh: nlst = (line.rstrip()).split() lst = lst + nlst # 利用有序字典去重,保留元素首次出现的顺序 lst = list(dict.fromkeys(lst)) print(lst)
方案3:手动调整索引(适合理解底层逻辑)
如果坚持用索引遍历,需在删除元素后调整索引值,避免跳过元素或索引越界:
fname = input("Enter file name: ") fh = open("romeo.txt") lst = list() for line in fh: nlst = (line.rstrip()).split() lst = lst + nlst i = 0 while i < len(lst): current_word = lst[i] if lst.count(current_word) > 1: while lst.count(current_word) > 1: print(current_word) lst.remove(current_word) else: i += 1 # 仅当当前元素无重复时,才递增索引 print(lst)
内容的提问来源于stack exchange,提问作者Youssef Wael
相关产品推荐
相关产品推荐

