Python中如何按不同字符序列拆分字符串生成对应元素列表?
字符串按指定序列拆分实现方案
核心拆分规则
拆分过程严格遵循要求:
- 从字符串起始位置从左到右遍历,全程不回溯、不重复读取字符
- 合法拆分元素固定为三类:
"P"(长度1)、"~O"(长度2)、"O~"(长度2) - 匹配优先级:优先匹配当前位置开头的2长度合法元素,匹配不到再匹配1长度的
"P",从根源避免错拆、漏拆
可直接运行的Python实现
def split_target_str(input_str: str) -> list[str]: valid_two_length_tokens = {"~O", "O~"} result = [] current_idx = 0 str_length = len(input_str) while current_idx < str_length: # 优先尝试匹配2长度的合法元素 if current_idx + 1 < str_length and input_str[current_idx:current_idx+2] in valid_two_length_tokens: result.append(input_str[current_idx:current_idx+2]) current_idx += 2 # 2长度匹配失败则匹配单字符P elif input_str[current_idx] == "P": result.append("P") current_idx += 1 # 题目给定的输入场景不会触发该异常分支 else: raise ValueError(f"索引{current_idx}位置存在无法匹配的非法字符: {input_str[current_idx]}") return result # 测试示例输入 a = "~O~O~O~OP" b = "PO~O~~OO~" c = "~O~O~O~OP~O~OO~" list_a = split_target_str(a) list_b = split_target_str(b) list_c = split_target_str(c) # 输出结果和预期完全一致 print(list_a) # ['~O', '~O', '~O', '~O', 'P'] print(list_b) # ['P', 'O~', 'O~', '~O', 'O~'] print(list_c) # ['~O', '~O', '~O', '~O', 'P', '~O', '~O', 'O~']
逻辑校验
以输入b = "PO~O~~OO~"为例,逐位匹配流程如下:
- 索引0位置:前2位为
"PO"不属于合法2长度元素,匹配单字符"P"加入结果,指针跳到1 - 索引1位置:前2位为
"O~"属于合法2长度元素,加入结果,指针跳到3 - 索引3位置:前2位为
"O~"属于合法2长度元素,加入结果,指针跳到5 - 索引5位置:前2位为
"~O"属于合法2长度元素,加入结果,指针跳到7 - 索引7位置:前2位为
"O~"属于合法2长度元素,加入结果,指针跳到9,遍历结束,输出和预期完全匹配。
内容的提问来源于stack exchange,提问作者A.Gng
相关产品推荐
相关产品推荐

