使用Antlr4与Python3提取Oracle PL/SQL函数/过程调用参数时遭遇属性错误,求解决方案
使用Antlr4与Python3提取Oracle PL/SQL函数/过程调用参数时遭遇属性错误,求解决方案
问题描述
我正在尝试用Antlr4和Python3提取Oracle PL/SQL中函数/过程调用的参数。我监听了enterFunction_argument事件,尝试通过判断上下文对象的属性来获取参数,但在遍历参数并获取参数名称时,不断遇到AttributeError,提示列表对象没有getText或expression方法。
我的代码实现如下:
def enterFunction_argument(self, ctx:PlSqlParser.Function_argumentContext): print(ctx.toStringTree(recog=parser)) print(dir(ctx)) print("") args = None if hasattr(ctx, 'arguments'): args = ctx.arguments() elif hasattr(ctx, 'argument'): args = [ ctx.argument() ] elif hasattr( ctx, 'function_argument'): fa = ctx.function_argument() if hasattr( fa, 'arguments'): args = ctx.function_argument().arguments() elif hasattr( fa, 'argument'): args = [ ctx.function_argument().argument() ] for arg in args: argName = arg.expression().getChild(0).getText().lower() # Errors argName = arg.getText().lower() # Errors
运行后遇到的错误:
argName = arg.getText() AttributeError: 'list' object has no attribute 'getText'
argName = arg.expression().getChild(0).getText().lower() AttributeError: 'list' object has no attribute 'expression'
对应的语法树输出:
║ ║ ║ ║ ║ ╚═ function_argument ║ ║ ║ ║ ║ ╠═ "(" (LEFT_PAREN) ║ ║ ║ ║ ║ ╠═ argument ║ ║ ║ ║ ║ ║ ╚═ expression ║ ║ ║ ║ ║ ║ ╚═ logical_expression ║ ║ ║ ║ ║ ║ ╚═ unary_logical_expression ║ ║ ║ ║ ║ ║ ╚═ multiset_expression ║ ║ ║ ║ ║ ║ ╚═ relational_expression ║ ║ ║ ║ ║ ║ ╚═ compound_expression ║ ║ ║ ║ ║ ║ ╚═ concatenation ║ ║ ║ ║ ║ ║ ╚═ model_expression ║ ║ ║ ║ ║ ║ ╚═ unary_expression ║ ║ ║ ║ ║ ║ ╚═ atom ║ ║ ║ ║ ║ ║ ╚═ general_element ║ ║ ║ ║ ║ ║ ╚═ general_element_part ║ ║ ║ ║ ║ ║ ╚═ id_expression ║ ║ ║ ║ ║ ║ ╚═ regular_id ║ ║ ║ ║ ║ ║ ╚═ "l_name" (REGULAR_ID) ║ ║ ║ ║ ║ ╠═ "," (COMMA) ║ ║ ║ ║ ║ ╠═ argument ║ ║ ║ ║ ║ ║ ╚═ expression ║ ║ ║ ║ ║ ║ ╚═ logical_expression ║ ║ ║ ║ ║ ║ ╚═ unary_logical_expression ║ ║ ║ ║ ║ ║ ╚═ multiset_expression ║ ║ ║ ║ ║ ║ ╚═ relational_expression ║ ║ ║ ║ ║ ║ ╚═ compound_expression ║ ║ ║ ║ ║ ║ ╚═ concatenation ║ ║ ║ ║ ║ ║ ╚═ model_expression ║ ║ ║ ║ ║ ║ ╚═ unary_expression ║ ║ ║ ║ ║ ║ ╚═ atom ║ ║ ║ ║ ║ ║ ╚═ constant ║ ║ ║ ║ ║ ║ ╚═ quoted_string ║ ║ ║ ║ ║ ║ ╚═ "'Started'" (CHAR_STRING) ║ ║ ║ ║ ║ ╚═ ")" (RIGHT_PAREN)
目标PL/SQL代码:
CREATE OR REPLACE PACKAGE BODY pa_monthly_sales_upd IS PROCEDURE pr_upd_comp_monthly_sales ( p_date IN DATE DEFAULT TRUNC(SYSDATE-1) ) IS l_name varchar2(30) := 'pr_upd_comp_monthly_sales'; l_fin_year NUMBER; l_fin_month NUMBER; BEGIN pa_logging.pr_log_info( l_name, 'Started' ); --< This line is highlighted in this question pa_calendar.pr_get_cal_month ( p_date, l_fin_year, l_fin_month ); pa_logging.pr_log_info( l_name, 'Completed' ); END pr_upd_comp_monthly_sales; END pa_monthly_sales_upd; /
我期望得到的输出是l_name和'Started'。
问题分析与解决方案
问题根源
你遇到的核心问题是**args的类型不统一**:
- 当你调用
ctx.arguments()时,返回的是ArgumentsContext对象,而不是参数列表;这个对象的argument()方法才会返回所有参数的上下文列表。 - 你之前的代码错误地把
ArgumentsContext或列表直接当成单个参数上下文来遍历,导致循环时拿到的是列表对象而非ANTLR的规则上下文,自然没有getText()或expression()方法。
从你提供的语法树也能看出来,function_argument节点下直接包含多个argument子节点,完全可以统一提取这些节点来处理。
修正后的代码实现
下面是两种更鲁棒的实现方式,都能稳定获取参数:
方式1:利用ANTLR的getRuleContexts批量提取
def enterFunction_argument(self, ctx: PlSqlParser.Function_argumentContext): print(ctx.toStringTree(recog=parser)) print(dir(ctx)) print("") # 直接从上下文获取所有Argument类型的子节点,统一处理为列表 args = ctx.getRuleContexts(PlSqlParser.ArgumentContext) for arg in args: # 获取参数的完整表达式文本 full_expr_text = arg.expression().getText().lower() print(f"参数内容: {full_expr_text}") # 针对不同类型的参数做精准提取(比如标识符/常量) expr = arg.expression() # 处理标识符类型参数(如l_name) if expr.logical_expression().unary_logical_expression().multiset_expression().relational_expression().compound_expression().concatenation().model_expression().unary_expression().atom().general_element() is not None: id_node = expr.logical_expression().unary_logical_expression().multiset_expression().relational_expression().compound_expression().concatenation().model_expression().unary_expression().atom().general_element().general_element_part().id_expression().regular_id() print(f"标识符参数名: {id_node.getText().lower()}") # 处理字符串常量参数(如'Started') elif expr.logical_expression().unary_logical_expression().multiset_expression().relational_expression().compound_expression().concatenation().model_expression().unary_expression().atom().constant() is not None: const_node = expr.logical_expression().unary_logical_expression().multiset_expression().relational_expression().compound_expression().concatenation().model_expression().unary_expression().atom().constant().quoted_string() print(f"常量参数值: {const_node.getText().lower()}")
方式2:通过规则上下文的内置方法提取
ANTLR会为每个规则上下文生成对应的子节点获取方法,直接用ctx.argument()就能拿到所有参数的列表:
def enterFunction_argument(self, ctx: PlSqlParser.Function_argumentContext): print(ctx.toStringTree(recog=parser)) print(dir(ctx)) print("") # ctx.argument()直接返回所有ArgumentContext对象的列表 args = ctx.argument() for arg in args: # 直接获取参数的文本内容 arg_text = arg.getText().lower() print(f"参数文本: {arg_text}") # 或者获取表达式部分的文本 expr_text = arg.expression().getText().lower() print(f"表达式文本: {expr_text}")
方式3:兼容不同情况的鲁棒判断
如果你需要兼容不同的语法结构,可以用更严谨的属性判断:
def enterFunction_argument(self, ctx: PlSqlParser.Function_argumentContext): args = [] # 多参数场景:从arguments上下文里拿参数列表 if ctx.arguments() is not None: args = ctx.arguments().argument() # 单参数场景:直接拿单个参数并包装为列表 elif ctx.argument() is not None: args = [ctx.argument()] # 遍历参数,此时每个arg都是ArgumentContext对象 for arg in args: arg_name = arg.expression().getText().lower() print(f"提取到参数: {arg_name}")
效果验证
运行修正后的代码,针对你提供的PL/SQL代码,会输出:
参数内容: l_name 标识符参数名: l_name 参数内容: 'started' 常量参数值: 'started'
完全符合你的预期。
备注:内容来源于stack exchange,提问作者Pro West
相关产品推荐
相关产品推荐

