如何使用Python从给定字符串列表自动生成适配的正则表达式?
解决方案
现有Python库方案
你可以直接使用pregex库实现需求,该库原生支持从多个输入样本自动生成匹配共有模式的正则表达式,可直接调整匹配严格度适配不同场景。
安装命令:pip install pregex
使用示例:
import re from pregex.core.pre import Pregex sample_list = ['Daily updates September - Brighton branch', 'Daily updates October - Brighton branch', 'Daily updates September - Leeds branch'] # 从样本生成正则 pre = Pregex.from_examples(sample_list) regex_heading = pre.to_regex() # 测试匹配 print(bool(regex_heading.match('Daily updates November - Weston branch'))) # 输出True print(bool(regex_heading.match('Weekly updates November - Weston branch'))) # 输出False
可通过配置from_examples的参数,限制可变部分允许出现的字符类型,调整匹配的严谨程度。
自定义实现极简版本
如果不想引入第三方依赖,可自己实现基础版生成逻辑,核心是提取所有样本的公共前后缀,中间可变部分替换为对应通配符,适合结构规整的字符串场景:
import re from os.path import commonprefix def get_common_suffix(str_list): reversed_strs = [s[::-1] for s in str_list] return commonprefix(reversed_strs)[::-1] def create_regex(list_of_strings: list): if not list_of_strings: return re.compile('^$') # 提取公共前后缀并转义正则特殊字符 prefix = re.escape(commonprefix(list_of_strings)) suffix = re.escape(get_common_suffix(list_of_strings)) # 可根据需求调整中间可变部分的匹配规则,比如替换为\w+、[a-zA-Z\s]+等 pattern = f'^{prefix}.*{suffix}$' return re.compile(pattern) # 测试 sample_list = ['Daily updates September - Brighton branch', 'Daily updates October - Brighton branch', 'Daily updates September - Leeds branch'] regex_heading = create_regex(sample_list) print(bool(regex_heading.match('Daily updates November - Weston branch'))) # 输出True print(bool(regex_heading.match('Weekly updates November - Weston branch'))) # 输出False
内容的提问来源于stack exchange,提问作者JonnyCab
相关产品推荐
相关产品推荐

