You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何定位字符串中子串多次出现的索引并修改Python函数适配?

问题1:获取字符串中子串多次出现的位置索引列表

咱们先明确需求:你要的是子串作为独立单词的单词索引(按空格分割后的列表位置),还是子串在原字符串里的字符位置索引?两种情况的实现方式不一样,我都给你列出来:

情况1:获取子串作为独立单词的索引

比如原字符串是"hello world hello python hello",子串"hello"对应的单词索引是0、2、4。代码可以这么写:

def find_substring_word_indices(s, substring):
    words = s.split(" ")
    indices = []
    for idx, word in enumerate(words):
        if word == substring:
            indices.append(str(idx))
    return " ".join(indices)

# 用例测试
s = "hello world hello python hello"
sub = "hello"
print(find_substring_word_indices(s, sub))  # 输出: "0 2 4"

情况2:获取子串在原字符串中的字符位置

比如"ababa"里找"aba",它的起始字符位置是0和2。这里用str.find()循环搜索,每次找到后从下一个位置继续:

def find_substring_char_indices(s, substring):
    indices = []
    start_pos = 0
    sub_length = len(substring)
    while start_pos <= len(s) - sub_length:
        current_pos = s.find(substring, start_pos)
        if current_pos == -1:
            break
        indices.append(str(current_pos))
        # 如果允许子串重叠匹配(比如"aaaa"找"aa"要返回0、1、2),就用start_pos = current_pos + 1
        # 如果不允许重叠,就用start_pos = current_pos + sub_length
        start_pos = current_pos + 1
    return " ".join(indices)

# 用例测试
s = "ababa"
sub = "aba"
print(find_substring_char_indices(s, sub))  # 输出: "0 2"

问题2:修改函数识别同一书名的多个实例

原函数的问题很明确:msg_split.index(last_item)只会返回第一个匹配项的索引,所以同一书名出现多次时,后面的实例就被忽略了。要解决这个问题,咱们得遍历msg_split的所有元素,找到所有匹配last_item的位置,同时还要保证不修改msg_split——这好办,提前把它存起来就好。

修改后的函数如下,我还加了个小检查避免误判(比如单独出现的章节号被当成书名的一部分):

def get_books(msg):
    results = []
    msg_split = msg.split(" ")  # 只分割一次,全程不修改这个列表
    for key, value in books.items():
        for item in value:
            if item in msg:
                last_item = item.split(" ")[-1]
                # 遍历msg_split的每一个索引,找所有匹配last_item的位置
                for idx, word in enumerate(msg_split):
                    if word == last_item:
                        # 额外检查:确保msg中对应位置的单词拼接起来等于完整的item
                        # 比如item是"Exodus 1:1",要确认从idx往前数1个位置的单词是"Exodus"
                        item_word_count = len(item.split(" "))
                        start_idx = idx - (item_word_count - 1)
                        if start_idx >= 0:
                            matched_segment = " ".join(msg_split[start_idx:idx+1])
                            if matched_segment == item:
                                results.append((key, idx))
    # 按索引排序,保证结果顺序和书名出现顺序一致
    results.sort(key=lambda x: x[1])
    return results

关键修改点:

  1. 提前把msg.split(" ")赋值给msg_split,只做一次分割,完全符合“不修改msg_split”的要求;
  2. 把msg_split.index(last_item)换成enumerate(msg_split)遍历,收集所有匹配的索引;
  3. 增加了匹配验证:避免单独的章节号(比如"1:1")被误判成书名的一部分,确保拼接后的片段和完整的item一致;
  4. 最后给结果按索引排序,返回的顺序和书名在msg里出现的顺序一致。

测试用例:

假设你的books字典是这样的:

books = {
    'exod': ['Exodus 1:1', 'Exodus 1:2'],
    'john': ['John 1:2']
}

输入"Exodus 1:1 blah blah blah Exodus 1:2",修改后的函数会返回[('exod', 0), ('exod', 5)],完美解决你的问题~


内容的提问来源于stack exchange,提问作者user2577858

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:26:01