如何查看AWS Glue ETL作业中的print语句输出?
Create and join tables
customer_churn = glueContext.create_dynamic_frame.from_catalog(database=db_name, table_name=tbl_customer_churn)
customer_churn = cust_joined.join(paths1=["customer id"], paths2=['id'], frame2=other_table)
logger.info(f"Customer_churn_joined:\n")
customer_churn.printSchema()
---- Write out the combined file ----
s_customer_churn = customer_churn.toDF().select("customer id")
logger.info(f"Customer_churn_just_cust_id:\n")
s_customer_churn.printSchema()
s_customer_churn.write.option("header","true").format("csv").mode('Overwrite').save(output_dir)
logger.info("output_dir:" + output_dir)
#### 其他相关信息: - Glue版本:5 - 类型:Spark - 语言:Python3 - 作业可观测性指标:已开启 - 持续日志记录:已开启 ### 已尝试操作 我查看了持续日志标签,能看到logger语句,但没有print语句输出。我看到输出日志会发送到CloudWatch(截图右下角),点击该链接后,所有日志中均未找到我的print语句。为何查找起来如此困难?  --- ### 回答 #### 1. print语句的输出位置 在Glue Spark作业中,`print()`或`printSchema()`的输出不会进入默认的持续日志流,而是被发送到**CloudWatch的`/aws-glue/jobs/output`日志组**里。这个日志组对应作业驱动程序和执行器的标准输出/错误日志,你需要在这里查找print内容。 日志流的命名格式通常是`job-<你的作业名>-run-<运行ID>-<driver/executor标识>`,直接搜索作业名就能快速定位。 #### 2. 替代print的更便捷调试方法 既然`printSchema()`返回`None`,可以直接把Schema序列化为字符串写入logger,这样就能在你已经能看到的日志里查看: ```python # 处理DynamicFrame的Schema schema_json = customer_churn.schema().json() logger.info(f"Customer_churn_joined Schema:\n{schema_json}") # 处理DataFrame的Schema schema_json = s_customer_churn.schema.json() logger.info(f"Customer_churn_just_cust_id Schema:\n{schema_json}")
3. 确保日志配置正确
- 检查作业配置的“日志选项”,确认已开启驱动程序日志和执行器日志的收集,这两个选项默认可能未启用。
- 把日志级别设置为
INFO或DEBUG,避免print输出被日志过滤规则拦截。
内容的提问来源于stack exchange,提问作者d-gg
相关产品推荐
相关产品推荐

