导入Jupyter Notebook函数时出现NameError: name 'athena' is not defined错误
问题原因及解决方案
原因分析
query_distinct_data()单独运行正常,是因为所在的map_distinct_data笔记本全局作用域里已经定义了athena和result_output_location变量。但导入到主笔记本后,函数的作用域仍隶属于map_distinct_data模块,该模块未定义这两个变量,主笔记本的全局变量不会自动注入到导入函数的作用域中,因此触发NameError。
其他导入的涉及Athena的函数能正常运行,大概率是因为它们要么在函数内部创建了Athena客户端,要么通过参数接收了所需的客户端/配置变量。
解决方案
推荐采用参数传递的方式解耦依赖,这是最规范且易维护的做法:
1. 修改map_distinct_data中的函数
将athena客户端和result_output_location作为参数传入函数:
def query_distinct_data(athena, result_output_location): query = "SELECT DISTINCT * from fire_data.rfs_fire_data where state in ('NSW','VIC','QLD')" response = athena.start_query_execution( QueryString=query, ResultConfiguration={"OutputLocation": result_output_location}) return response["QueryExecutionId"]
2. 主笔记本中调用时传入参数
在主笔记本调用函数时,把已定义的athena和result_output_location传递进去:
execution_id = query_distinct_data(athena, result_output_location)
可选方案(封装配置模块)
如果多个模块都需要用到Athena客户端和配置,可以单独封装一个配置模块,避免重复代码:
- 创建
config.ipynb笔记本,内容如下:
import boto3 aws_region = "ap-southeast-2" result_output_location = "s3://camgoo2-rfs-visualisation/query_results/" athena = boto3.client("athena", region_name=aws_region)
- 在
map_distinct_data中导入配置:
from ipynb.fs.full.config import athena, result_output_location def query_distinct_data(): query = "SELECT DISTINCT * from fire_data.rfs_fire_data where state in ('NSW','VIC','QLD')" response = athena.start_query_execution( QueryString=query, ResultConfiguration={"OutputLocation": result_output_location}) return response["QueryExecutionId"]
- 主笔记本也从
config导入变量,移除重复定义的内容:
from ipynb.fs.full.config import aws_region, result_output_location, bucket, athena
内容的提问来源于stack exchange,提问作者camgoo2
相关产品推荐
相关产品推荐

