能否在Foundry code authoring中获取dataset最后sync日期并作为列导入
Palantir Foundry 数据集同步时间获取与写入方案
你提出的两个需求都可以在Foundry Code Authoring环境中实现,具体操作方法如下:
- 获取数据集最后同步日期
可以直接通过Transforms框架提供的元数据接口读取对应数据集的最后同步时间,返回结果为带UTC时区的datetime类型,可直接用于时间范围判断。核心调用方法为:
from transforms.api import Input # 此处替换为你的源数据集路径 source = Input("/组织/项目/源数据集路径") last_sync_time = source.get_metadata().last_modified_time
- 将同步日期作为列导入目标数据集
你可以通过PySpark的lit函数将获取到的同步时间作为常量列添加到目标数据集中,完整代码示例如下:
from transforms.api import transform, Input, Output from pyspark.sql.functions import lit from datetime import datetime, timezone @transform( output=Output("/组织/项目/目标数据集路径"), source=Input("/组织/项目/源数据集路径") ) def compute(source, output): # 1. 获取源数据集最后同步时间 last_sync_dt = source.get_metadata().last_modified_time # 可选:判断同步时间是否在指定范围内,不符合则终止任务 start_time = datetime(2024, 1, 1, tzinfo=timezone.utc) end_time = datetime(2024, 6, 30, tzinfo=timezone.utc) if not (start_time <= last_sync_dt <= end_time): raise ValueError(f"源数据集最后同步时间{last_sync_dt}不在指定的[2024-01-01, 2024-06-30]范围内") # 2. 读取源数据并添加同步时间列 source_df = source.dataframe() result_df = source_df.withColumn("source_last_sync_time", lit(last_sync_dt)) # 3. 写入目标数据集 output.write_dataframe(result_df)
内容的提问来源于stack exchange,提问作者Chloe Lathe
相关产品推荐
相关产品推荐

