You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修复基于索引列表为字符串标记添加下划线的代码问题

问题背景

我们希望通过索引列表为字符串中的指定标记添加下划线,预期输出格式为:'7 Waitohu Road _York_Bay Co Manager _York_Bay Asst Co Dir _Central_Lower_Hutt General Hand _Wainuiomata School Caretaker'。

当前代码及问题

现有Python代码如下:

# 说明:
# 遍历拆分后的列表直到到达指定索引位置。如果当前索引后紧跟另一个索引,就给对应的拆分标记添加无空格的下划线
# 预期输出应为:'7 Waitohu Road _York_Bay Co Manager _York_Bay Asst Co Dir _Central_Lower_Hutt General Hand _Wainuiomata School Caretaker' 

# 包含郊区名称及其在字符串中索引位置的列表
uniqueList = ['York', 3, 'Bay', 4, 'York', 7, 'Bay', 8, 'Central', 12, 'Lower', 13, 'Hutt', 14, 'Wainuiomata', 17]

# 提取列表中的索引部分
indexes = [3, 4, 7, 8, 12, 13, 14, 17]

# 示例字符串
line = '7 Waitohu Road York Bay Co Manager York Bay Asst Co Dir Central Lower Hutt General Hand Wainuiomata School Caretaker'

# 将字符串拆分为单词列表以便按索引处理
splits = line.split(' ')

# 遍历索引
for i in range(len(indexes)):
    check = indexes[i]
    for j in range(len(splits)):
        if j == check and (i + 1 < len(indexes)):
            # 判断下一个索引是否是当前索引+1
            next_idx = indexes[i + 1]
            if 1 == next_idx - check:
                splits[j] = '_' + splits[j] + '_' + splits[j + 1]            
        else:
            if j == check:
                splits[j] = '_' + splits[j]

# 生成结果
newLine = ' '.join(splits)
print(newLine)

当前代码运行输出为:

7 Waitohu Road _York_Bay Bay Co Manager _York_Bay Bay Asst Co Dir _Central_Lower _Lower_Hutt Hutt General Hand _Wainuiomata School Caretaker

需要解决的问题

  • 移除输出中的重复单词(如Bay、Hutt);
  • 实现多词连续下划线拼接,得到_Central_Lower_Hutt格式的结果。
解决方案

原代码的问题在于仅修改了当前索引的单词,未处理后续重复项,且连续多词的拼接逻辑有误。正确思路是先将连续索引分组,再对每组单词进行下划线拼接,同时标记已处理索引避免重复。

修改后的代码如下:

# 包含郊区名称及其在字符串中索引位置的列表
uniqueList = ['York', 3, 'Bay', 4, 'York', 7, 'Bay', 8, 'Central', 12, 'Lower', 13, 'Hutt', 14, 'Wainuiomata', 17]

# 提取索引部分
indexes = [3, 4, 7, 8, 12, 13, 14, 17]

# 示例字符串
line = '7 Waitohu Road York Bay Co Manager York Bay Asst Co Dir Central Lower Hutt General Hand Wainuiomata School Caretaker'

splits = line.split(' ')
processed = set()  # 记录已处理的索引,避免重复输出

# 将连续递增的索引分组,比如[3,4]、[7,8]、[12,13,14]、[17]
groups = []
current_group = [indexes[0]]
for idx in indexes[1:]:
    if idx == current_group[-1] + 1:
        current_group.append(idx)
    else:
        groups.append(current_group)
        current_group = [idx]
groups.append(current_group)

# 处理每个分组
for group in groups:
    # 把分组内的单词用下划线连接,开头加一个下划线
    joined_str = '_' + '_'.join(splits[i] for i in group)
    # 替换分组第一个位置的单词为拼接后的字符串
    splits[group[0]] = joined_str
    # 标记分组内后续索引为已处理
    for i in group[1:]:
        processed.add(i)

# 生成结果时跳过已处理的索引
newLine = ' '.join(word for idx, word in enumerate(splits) if idx not in processed)
print(newLine)

代码说明

  1. 索引分组:将连续递增的索引归类为一组,解决多词连续拼接的需求;
  2. 拼接替换:对每组内的单词进行下划线拼接,替换分组首个位置的单词;
  3. 去重处理:用集合记录被合并的索引,最终生成结果时跳过这些索引,消除重复单词。

运行后输出符合预期:

7 Waitohu Road _York_Bay Co Manager _York_Bay Asst Co Dir _Central_Lower_Hutt General Hand _Wainuiomata School Caretaker

内容的提问来源于stack exchange,提问作者Dave

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 13:33:16