如何从指定字符串中提取NOT x关联的数值?
提取被NOT修饰的变量x的赋值数值
需要从以下字符串中提取被NOT修饰的变量x的赋值数值:
(req.idf=6ca9a AND (req.ster=201 OR req.ster=st_home) AND (req.ste=hi OR req.ster=hijst_iuer OR ((req.ster=laHome OR req.ster=laHome_Jtre) AND (tax=IN OR taxIP=MX))) AND NOT x=3422 AND (NOT (x=u259 OR x=1132 OR x=1144))AND NOT (x=28743 or x=09323 or x=12323) AND (x=113323 OR x=90323 OR x=1112123)
期望输出:[3422, 1132, 1144, 28743, 09323, 12323]
尝试过的正则表达式
仅能提取3422,无法覆盖NOT后接括号、括号内多个x赋值的场景:
(?:(?:NOT\s*x=)|(?:NOTx=))(\d+)
测试场景补充
对于字符串:
x= 3424 AND NOT (x=34343 OR x=1214 OR x=11121 AND (x=23232 OR x=1121))
应提取数值:[34343, 1214, 11121, 23232, 1121]
核心需求
- 检测到
NOT后,若后续为x=或x=,提取x对应的数值; - 检测到
NOT后,若后续为(或(,则在匹配的括号范围内,提取所有x对应的数值。
当前代码及问题
自己编写的Python代码如下,仅返回[3422],无法满足需求:
def extract_not_x(target): result = [] i = 0 while i < len(target): if target[i:i+3] == "NOT": i += 3 if i < len(target) and (target[i:i+2] == "x=" or target[i:i+3] == " x="): i += 2 if target[i+1] == "=" else 3 j = i while j < len(target) and target[j].isdigit(): j += 1 result.append(int(target[i:j])) i = j elif i < len(target) and (target[i:i+2] == "( " or target[i] == "("): i += 1 while i < len(target) and target[i] != ")": if target[i:i+2] == "x=" or target[i:i+3] == " x=": i += 2 if target[i+1] == "=" else 3 j = i while j < len(target) and target[j].isdigit(): j += 1 result.append(int(target[i:j])) i = j else: i += 1 i += 1 else: i += 1 return result print(extract_not_x("(req.idf=6ca9a AND (req.ster=201 OR req.ster=st_home) AND (req.ste=hi OR req.ster=hijst_iuer OR ((req.ster=laHome OR req.ster=laHome_Jtre) AND (tax=IN OR taxIP=MX))) AND NOT x=3422 AND (NOT (x=u259 OR x=1132 OR x=1144))AND NOT (x=28743 or x=09323 or x=12323) AND (x=113323 OR x=90323 OR x=1112123)"))
解决方案1:使用支持递归的正则表达式
Python的re模块支持递归正则,可以匹配嵌套括号。先匹配NOT后的所有目标内容,再从中提取x的数值:
import re def extract_not_x(target): result = [] # 匹配NOT后两种情况:直接x=数字 或 嵌套括号内容 not_pattern = re.compile(r'NOT\s*(?:x=(\d+)|(\((?:[^()]|(?2))*\)))', re.IGNORECASE) # 从任意内容中提取x=后的数字 x_pattern = re.compile(r'x=(\d+)', re.IGNORECASE) for match in not_pattern.finditer(target): if match.group(1): result.append(match.group(1)) elif match.group(2): # 提取括号内所有x=的数字 for x_match in x_pattern.finditer(match.group(2)): result.append(x_match.group(1)) return result # 测试示例 test_str1 = "(req.idf=6ca9a AND (req.ster=201 OR req.ster=st_home) AND (req.ste=hi OR req.ster=hijst_iuer OR ((req.ster=laHome OR req.ster=laHome_Jtre) AND (tax=IN OR taxIP=MX))) AND NOT x=3422 AND (NOT (x=u259 OR x=1132 OR x=1144))AND NOT (x=28743 or x=09323 or x=12323) AND (x=113323 OR x=90323 OR x=1112123)" print(extract_not_x(test_str1)) # 输出 ['3422', '1132', '1144', '28743', '09323', '12323'] test_str2 = "x= 3424 AND NOT (x=34343 OR x=1214 OR x=11121 AND (x=23232 OR x=1121))" print(extract_not_x(test_str2)) # 输出 ['34343', '1214', '11121', '23232', '1121']
解决方案2:改进原始代码处理嵌套括号
修复原始代码中未处理嵌套括号、未跳过空格、未过滤非数字x赋值的问题:
def extract_not_x(target): result = [] i = 0 len_target = len(target) while i < len_target: if target[i:i+3] == "NOT": i += 3 # 跳过NOT后的所有空格 while i < len_target and target[i].isspace(): i += 1 if i < len_target and target[i:i+2].lower() == "x=": # 处理直接x=的情况 i += 2 while i < len_target and target[i].isspace(): i += 1 # 提取数字(保留0开头的字符串) j = i while j < len_target and target[j].isdigit(): j += 1 if j > i: result.append(target[i:j]) i = j elif i < len_target and target[i] == "(": # 处理嵌套括号,计算括号层级 i += 1 bracket_level = 1 start = i while i < len_target and bracket_level > 0: if target[i] == "(": bracket_level += 1 elif target[i] == ")": bracket_level -= 1 i += 1 # 提取括号内内容并查找x=的数字 bracket_content = target[start:i-1] k = 0 len_bracket = len(bracket_content) while k < len_bracket: if bracket_content[k:k+2].lower() == "x=": k += 2 while k < len_bracket and bracket_content[k].isspace(): k += 1 m = k while m < len_bracket and bracket_content[m].isdigit(): m += 1 if m > k: result.append(bracket_content[k:m]) k = m else: k += 1 else: i += 1 return result # 测试示例 test_str1 = "(req.idf=6ca9a AND (req.ster=201 OR req.ster=st_home) AND (req.ste=hi OR req.ster=hijst_iuer OR ((req.ster=laHome OR req.ster=laHome_Jtre) AND (tax=IN OR taxIP=MX))) AND NOT x=3422 AND (NOT (x=u259 OR x=1132 OR x=1144))AND NOT (x=28743 or x=09323 or x=12323) AND (x=113323 OR x=90323 OR x=1112123)" print(extract_not_x(test_str1)) # 输出 ['3422', '1132', '1144', '28743', '09323', '12323'] test_str2 = "x= 3424 AND NOT (x=34343 OR x=1214 OR x=11121 AND (x=23232 OR x=1121))" print(extract_not_x(test_str2)) # 输出 ['34343', '1214', '11121', '23232', '1121']
说明
- 正则方案更简洁,利用递归组处理嵌套括号,适合规则明确的场景;
- 代码方案手动处理括号层级,灵活性更高,适合需要自定义逻辑的场景;
- 两个方案均保留0开头的数值字符串,如需转为整数,在添加结果时用
int()转换即可。
内容的提问来源于stack exchange,提问作者shubham0001
相关产品推荐
相关产品推荐

