如何修复基于索引列表为字符串标记添加下划线的代码问题
问题背景
我们希望通过索引列表为字符串中的指定标记添加下划线,预期输出格式为:'7 Waitohu Road _York_Bay Co Manager _York_Bay Asst Co Dir _Central_Lower_Hutt General Hand _Wainuiomata School Caretaker'。
当前代码及问题
现有Python代码如下:
# 说明: # 遍历拆分后的列表直到到达指定索引位置。如果当前索引后紧跟另一个索引,就给对应的拆分标记添加无空格的下划线 # 预期输出应为:'7 Waitohu Road _York_Bay Co Manager _York_Bay Asst Co Dir _Central_Lower_Hutt General Hand _Wainuiomata School Caretaker' # 包含郊区名称及其在字符串中索引位置的列表 uniqueList = ['York', 3, 'Bay', 4, 'York', 7, 'Bay', 8, 'Central', 12, 'Lower', 13, 'Hutt', 14, 'Wainuiomata', 17] # 提取列表中的索引部分 indexes = [3, 4, 7, 8, 12, 13, 14, 17] # 示例字符串 line = '7 Waitohu Road York Bay Co Manager York Bay Asst Co Dir Central Lower Hutt General Hand Wainuiomata School Caretaker' # 将字符串拆分为单词列表以便按索引处理 splits = line.split(' ') # 遍历索引 for i in range(len(indexes)): check = indexes[i] for j in range(len(splits)): if j == check and (i + 1 < len(indexes)): # 判断下一个索引是否是当前索引+1 next_idx = indexes[i + 1] if 1 == next_idx - check: splits[j] = '_' + splits[j] + '_' + splits[j + 1] else: if j == check: splits[j] = '_' + splits[j] # 生成结果 newLine = ' '.join(splits) print(newLine)
当前代码运行输出为:
7 Waitohu Road _York_Bay Bay Co Manager _York_Bay Bay Asst Co Dir _Central_Lower _Lower_Hutt Hutt General Hand _Wainuiomata School Caretaker
需要解决的问题
- 移除输出中的重复单词(如
Bay、Hutt); - 实现多词连续下划线拼接,得到
_Central_Lower_Hutt格式的结果。
解决方案
原代码的问题在于仅修改了当前索引的单词,未处理后续重复项,且连续多词的拼接逻辑有误。正确思路是先将连续索引分组,再对每组单词进行下划线拼接,同时标记已处理索引避免重复。
修改后的代码如下:
# 包含郊区名称及其在字符串中索引位置的列表 uniqueList = ['York', 3, 'Bay', 4, 'York', 7, 'Bay', 8, 'Central', 12, 'Lower', 13, 'Hutt', 14, 'Wainuiomata', 17] # 提取索引部分 indexes = [3, 4, 7, 8, 12, 13, 14, 17] # 示例字符串 line = '7 Waitohu Road York Bay Co Manager York Bay Asst Co Dir Central Lower Hutt General Hand Wainuiomata School Caretaker' splits = line.split(' ') processed = set() # 记录已处理的索引,避免重复输出 # 将连续递增的索引分组,比如[3,4]、[7,8]、[12,13,14]、[17] groups = [] current_group = [indexes[0]] for idx in indexes[1:]: if idx == current_group[-1] + 1: current_group.append(idx) else: groups.append(current_group) current_group = [idx] groups.append(current_group) # 处理每个分组 for group in groups: # 把分组内的单词用下划线连接,开头加一个下划线 joined_str = '_' + '_'.join(splits[i] for i in group) # 替换分组第一个位置的单词为拼接后的字符串 splits[group[0]] = joined_str # 标记分组内后续索引为已处理 for i in group[1:]: processed.add(i) # 生成结果时跳过已处理的索引 newLine = ' '.join(word for idx, word in enumerate(splits) if idx not in processed) print(newLine)
代码说明
- 索引分组:将连续递增的索引归类为一组,解决多词连续拼接的需求;
- 拼接替换:对每组内的单词进行下划线拼接,替换分组首个位置的单词;
- 去重处理:用集合记录被合并的索引,最终生成结果时跳过这些索引,消除重复单词。
运行后输出符合预期:
7 Waitohu Road _York_Bay Co Manager _York_Bay Asst Co Dir _Central_Lower_Hutt General Hand _Wainuiomata School Caretaker
内容的提问来源于stack exchange,提问作者Dave
相关产品推荐
相关产品推荐

