Python中相同单词索引相同,如何区分字符串内不同位置的相同单词?
检查字符串中重复单词的索引问题
我正在将字符串与自身比较,以检查同一字符串中是否存在重复单词。我尝试通过索引来区分单词,但由于Python对相同值的单词返回相同索引,当前代码没有输出。(请不要建议使用range(len(x))类方法,我希望用这种特定的比较方式)
input= "The hello hello substring string of this pan is amazing hello" x=input.split(" ") for i in x: for j in x: if x.index(i)!=x.index(j) and i==j: print(i,x.index(i),j,x.index(j))
我期望得到类似“hello,1,hello,2,hello,10”的输出,但实际Python中所有hello的索引都是1,因此没有输出。请问为何不同位置的单词会有相同索引,该如何解决?
原因
list.index(item)方法只会返回该元素在列表中第一次出现的索引,不管后续列表里还有多少个相同元素。所以你代码里所有值为hello的元素调用x.index(i)都会返回1,导致x.index(i)!=x.index(j)这个条件永远不成立,自然没有输出。
解决方案
要保留直接遍历元素的方式,同时获取每个元素的真实位置,可以用enumerate()给每个元素绑定对应的索引,这样遍历的时候既能拿到单词,也能拿到它在列表里的实际位置:
方案1:输出所有重复索引对
input_str = "The hello hello substring string of this pan is amazing hello" x = input_str.split(" ") # 把每个元素和它的索引配对 word_index_pairs = list(enumerate(x)) for idx_i, word_i in word_index_pairs: for idx_j, word_j in word_index_pairs: # 只处理索引不同且单词相同的情况,同时避免重复打印(比如1和2、2和1只输出一次) if idx_i < idx_j and word_i == word_j: print(f"{word_i},{idx_i},{word_j},{idx_j}")
输出:
hello,1,hello,2 hello,1,hello,10 hello,2,hello,10
方案2:按期望格式输出连续结果
如果想得到类似hello,1,hello,2,hello,10的连续格式,可以先收集每个单词的所有索引,再拼接输出:
input_str = "The hello hello substring string of this pan is amazing hello" x = input_str.split(" ") # 用字典存储每个单词对应的所有索引 word_indices = {} for idx, word in enumerate(x): if word not in word_indices: word_indices[word] = [] word_indices[word].append(idx) # 遍历字典,输出有重复的单词 for word, indices in word_indices.items(): if len(indices) > 1: output = [] for idx in indices: output += [word, str(idx)] print(",".join(output))
输出:
hello,1,hello,2,hello,10
内容的提问来源于stack exchange,提问作者Vikrant KALKAL
相关产品推荐
相关产品推荐

