如何用Python结合rules.list文件通过正则提取命名实体?
利用Python结合正则规则文件提取命名实体
步骤说明
- 解析规则文件:读取
rules.list中的每条规则,拆分出正则匹配模式和要提取的捕获组索引。比如规则app install (.*) 1,对应正则模式为app install (.*),需提取第1个捕获组的内容。 - 匹配命令字符串:遍历每条输入命令,用解析后的正则规则尝试匹配,匹配成功则提取对应捕获组的内容。
代码实现
import re def load_rules(rule_file): rules = [] with open(rule_file, 'r') as f: for line in f: line = line.strip() if not line: continue # 拆分规则:格式为 "匹配模式 捕获组索引" parts = line.rsplit(' ', 1) if len(parts) != 2: continue pattern_str, group_idx = parts try: group_idx = int(group_idx) # 编译正则,忽略大小写(可按需调整) pattern = re.compile(pattern_str, re.IGNORECASE) rules.append((pattern, group_idx)) except ValueError: # 跳过格式错误的规则 continue return rules def extract_entities(commands, rules): entities = [] for cmd in commands: matched = False for pattern, group_idx in rules: match = pattern.search(cmd) if match and group_idx <= len(match.groups()): entity = match.group(group_idx).strip() entities.append(entity) matched = True break # 匹配到第一条规则即停止,可按需修改 if not matched: entities.append(None) # 未匹配时的默认值,可按需调整 return entities # 示例调用 commands = [ "Open Youtube", "Install PlayStore App", "Go to Call of Duty app" ] # 需确保rules.list中的规则与命令对应,示例规则如下: # open (.*) 1 # install (.*) app 1 # go to (.*) app 1 rules = load_rules('rules.list') result = extract_entities(commands, rules) print(result) # 输出: ['Youtube', 'PlayStore', 'Call of Duty']
注意事项
- 保证
rules.list规则格式统一,每条规则末尾为捕获组索引数字 - 正则规则需与命令字符串格式匹配,比如
Open Youtube对应规则open (.*) 1,Install PlayStore App对应install (.*) app 1 - 可根据需求调整正则匹配选项,比如是否要求全匹配、是否区分大小写等
内容的提问来源于stack exchange,提问作者Turing101
相关产品推荐
相关产品推荐

