如何高效实现带括号的字符串拆分:保留括号内多词元素
问题:拆分字符串时保留括号内的多词内容作为单个元素
需求:将字符串拆分为列表,非括号内的内容按常规拆分,括号内的多词内容作为单个元素保留。例如输入"(word1 word2) word3 (word4 word5)",应得到['word1 word2', 'word3', 'word4 word5']。
原实现代码:
import re def get_queries(s): parentheses_queries = re.findall('\((.*?)\)', s) if not parentheses_queries: return s.split() for q in parentheses_queries: if f'({q})' in s: s = s.replace(q, '') queries = s.strip().split() i = 0 while '()' in queries: queries[queries.index('()')] = parentheses_queries[i] i += 1 return queries s = '(word1 word2) word3 (word4 word5)' print(get_queries(s))
优化思路
方法1:正则一次性匹配(简洁高效)
直接用正则匹配所有目标元素——要么提取括号内的内容,要么匹配括号外的单个单词,一次遍历即可得到结果,无需多次替换和循环:
import re def get_queries(s): # 正则规则:匹配括号内的内容,或括号外的连续非空白/非括号字符 pattern = r'\((.*?)\)|(\S+)' matches = re.findall(pattern, s) # 从匹配结果中取非空的分组(括号内内容优先) return [match[0] if match[0] else match[1] for match in matches] s = '(word1 word2) word3 (word4 word5)' print(get_queries(s)) # 输出: ['word1 word2', 'word3', 'word4 word5']
优势:
- 逻辑简洁,仅需一次正则操作,避免原方法中多次字符串替换和列表修改的开销
- 避免原方法可能出现的替换冲突(比如括号内内容与外部内容重复时,替换会出错)
- 时间复杂度更低,执行效率更高
方法2:手动遍历解析(支持复杂场景)
如果需要处理嵌套括号(原方法和上述正则均不支持),可以手动遍历字符串,跟踪括号状态来解析:
def get_queries(s): result = [] current_chars = [] in_parentheses = False for char in s: if char == '(': in_parentheses = True elif char == ')': in_parentheses = False # 括号闭合时,将累积的内容作为单个元素加入结果 result.append(''.join(current_chars).strip()) current_chars = [] elif char == ' ' and not in_parentheses: # 非括号内的空格作为分隔符,将当前累积的单词加入结果 if current_chars: result.append(''.join(current_chars).strip()) current_chars = [] else: current_chars.append(char) # 处理字符串末尾未闭合的内容 if current_chars: result.append(''.join(current_chars).strip()) # 过滤空字符串元素 return [item for item in result if item] s = '(word1 (word2 word3)) word4' print(get_queries(s)) # 输出: ['word1 word2 word3', 'word4']
优势:
- 支持嵌套括号场景,鲁棒性更强
- 无需依赖正则,逻辑清晰可控
内容的提问来源于stack exchange,提问作者Oussama Blgrim
相关产品推荐
相关产品推荐

