如何用正则表达式提取以But开头、typesetting前一词结尾的子串?
提取指定规则的子串解决方案
要提取以But开头、以typesetting前一词结尾的文本,可通过正则表达式的正向预查实现,既能避开typesetting本身,也能适配多字符串遍历的需求。
核心正则逻辑
正则表达式:But.*?(?=\stypesetting)
But:精准匹配固定起始字符串.*?:非贪婪模式匹配任意字符,确保只匹配到第一个typesetting之前的内容,避免过度匹配(?=\stypesetting):正向预查,匹配前面内容后紧跟空格+typesetting的位置,不会将typesetting纳入结果
代码示例(Python)
批量匹配所有符合规则的子串
import re # 示例字符串 res = "But also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing Lorem Ipsum passages, and more recently with desktop" # 遍历字符串列表时,循环处理每个字符串即可 results = re.findall(r'But.*?(?=\stypesetting)', res) print(results) # 输出: ['But also the leap into electronic']
提取单个匹配结果(优先取第一个)
如果只需要第一个符合规则的子串,用search效率更高:
match = re.search(r'But.*?(?=\stypesetting)', res) if match: target_substring = match.group() print(target_substring) # 输出: But also the leap into electronic
适配多字符串遍历场景
处理一组字符串时,将逻辑放入循环即可:
import re string_list = [ "But first test word typesetting ...", "But another example text typesetting ..." ] for s in string_list: match = re.search(r'But.*?(?=\stypesetting)', s) if match: print(match.group())
内容的提问来源于stack exchange,提问作者Munrock
相关产品推荐
相关产品推荐

