如何定位字符串中子串多次出现的索引并修改Python函数适配?
问题1:获取字符串中子串多次出现的位置索引列表
咱们先明确需求:你要的是子串作为独立单词的单词索引(按空格分割后的列表位置),还是子串在原字符串里的字符位置索引?两种情况的实现方式不一样,我都给你列出来:
情况1:获取子串作为独立单词的索引
比如原字符串是"hello world hello python hello",子串"hello"对应的单词索引是0、2、4。代码可以这么写:
def find_substring_word_indices(s, substring): words = s.split(" ") indices = [] for idx, word in enumerate(words): if word == substring: indices.append(str(idx)) return " ".join(indices) # 用例测试 s = "hello world hello python hello" sub = "hello" print(find_substring_word_indices(s, sub)) # 输出: "0 2 4"
情况2:获取子串在原字符串中的字符位置
比如"ababa"里找"aba",它的起始字符位置是0和2。这里用str.find()循环搜索,每次找到后从下一个位置继续:
def find_substring_char_indices(s, substring): indices = [] start_pos = 0 sub_length = len(substring) while start_pos <= len(s) - sub_length: current_pos = s.find(substring, start_pos) if current_pos == -1: break indices.append(str(current_pos)) # 如果允许子串重叠匹配(比如"aaaa"找"aa"要返回0、1、2),就用start_pos = current_pos + 1 # 如果不允许重叠,就用start_pos = current_pos + sub_length start_pos = current_pos + 1 return " ".join(indices) # 用例测试 s = "ababa" sub = "aba" print(find_substring_char_indices(s, sub)) # 输出: "0 2"
问题2:修改函数识别同一书名的多个实例
原函数的问题很明确:msg_split.index(last_item)只会返回第一个匹配项的索引,所以同一书名出现多次时,后面的实例就被忽略了。要解决这个问题,咱们得遍历msg_split的所有元素,找到所有匹配last_item的位置,同时还要保证不修改msg_split——这好办,提前把它存起来就好。
修改后的函数如下,我还加了个小检查避免误判(比如单独出现的章节号被当成书名的一部分):
def get_books(msg): results = [] msg_split = msg.split(" ") # 只分割一次,全程不修改这个列表 for key, value in books.items(): for item in value: if item in msg: last_item = item.split(" ")[-1] # 遍历msg_split的每一个索引,找所有匹配last_item的位置 for idx, word in enumerate(msg_split): if word == last_item: # 额外检查:确保msg中对应位置的单词拼接起来等于完整的item # 比如item是"Exodus 1:1",要确认从idx往前数1个位置的单词是"Exodus" item_word_count = len(item.split(" ")) start_idx = idx - (item_word_count - 1) if start_idx >= 0: matched_segment = " ".join(msg_split[start_idx:idx+1]) if matched_segment == item: results.append((key, idx)) # 按索引排序,保证结果顺序和书名出现顺序一致 results.sort(key=lambda x: x[1]) return results
关键修改点:
- 提前把
msg.split(" ")赋值给msg_split,只做一次分割,完全符合“不修改msg_split”的要求; - 把
msg_split.index(last_item)换成enumerate(msg_split)遍历,收集所有匹配的索引; - 增加了匹配验证:避免单独的章节号(比如"1:1")被误判成书名的一部分,确保拼接后的片段和完整的
item一致; - 最后给结果按索引排序,返回的顺序和书名在msg里出现的顺序一致。
测试用例:
假设你的books字典是这样的:
books = { 'exod': ['Exodus 1:1', 'Exodus 1:2'], 'john': ['John 1:2'] }
输入"Exodus 1:1 blah blah blah Exodus 1:2",修改后的函数会返回[('exod', 0), ('exod', 5)],完美解决你的问题~
内容的提问来源于stack exchange,提问作者user2577858
相关产品推荐
相关产品推荐

