如何优化正则实现按最大长度识词分割字符串且不含前导空格
无前置空格的按词分割正则方案
方案1:调整正则表达式
修改原正则,确保匹配的子串从非空格字符起始,避免前导空格:
(?<!\S)\w+(?:\s+\w+){0,}?(?<=.{20,})(?=\s)|(?<!\S).{1,20}(?=\s)|(?<!\S).+$
规则说明:
(?<!\S):限定匹配起始位置为字符串开头或空格之后,从根源避免前导空格\w+(?:\s+\w+){0,}?(?<=.{20,})(?=\s):匹配长度≥20字符的完整单词序列,仅在空格处截断(?<!\S).{1,20}(?=\s):匹配长度不足20字符、可在空格处截断的片段(?<!\S).+$:匹配最后一段剩余内容
方案2:匹配结果后处理
若不想修改正则,可直接对匹配结果做去空格处理,以Python代码为例:
import re input_text = "this is an input example of one sentence that contains a bit of words and must be split" split_pattern = re.compile(r'\b[\w\s]{20,}?(?=\s)|.+$') cleaned_result = [segment.strip() for segment in split_pattern.findall(input_text) if segment.strip()]
执行后得到的结果即为:
[ "this is an input example", "of one sentence that", "contains a bit of words", "and must be split" ]
内容的提问来源于stack exchange,提问作者OfirD
相关产品推荐
相关产品推荐

