如何在Palantir Foundry中使用Function按多属性分组聚合
Palantir Foundry 基于Function的多属性分组聚合实现
以下实现直接基于Foundry Ontology原生聚合能力实现,计算下推到存储层,避免全量数据拉取导致的内存、超时问题:
前置准备
- 确保你的Function已经关联Schedule对象类型的Ontology访问权限
- 确认Schedule各属性的API名称和描述一致:
date(日期类型)、shift_type(字符串类型)、department(字符串类型)、hours worked(数值类型)
完整实现代码
import { Function, Objects, LocalDate } from "@foundry/functions-api"; import { Schedule } from "@foundry/ontology-api"; export class ScheduleAggregation { /** * 按日期范围统计各日期、班次、部门的总工时 * @param startDate 统计开始日期 * @param endDate 统计结束日期 */ @Function() public async calculateTotalWorkHours( startDate: LocalDate, endDate: LocalDate ): Promise<Array<{ date: LocalDate, shiftType: string, department: string, totalHours: number }>> { // 基础日期合法性校验 if (startDate.isAfter(endDate)) { throw new Error("传入的开始日期不能晚于结束日期"); } // 过滤时间区间内的排班数据 const filteredSchedules = Objects.search() .schedule() .filter(item => item.date.gte(startDate).and(item.date.lte(endDate))); // 按三个维度分组,对工时求和 const aggResult = await filteredSchedules.groupBy(record => ({ date: record.date.exact(), shiftType: record.shift_type.exact(), department: record.department.exact() })) .aggregations(a => ({ // 注意属性名带空格,需要用中括号形式访问 totalHours: a["hours worked"].sum() })) .all(); // 整理为目标输出结构 return aggResult.map(item => ({ date: item.group.date, shiftType: item.group.shiftType, department: item.group.department, totalHours: item.aggregations.totalHours })); } }
关键注意事项
- 禁止直接调用
.all()拉取全量Schedule对象到Function内存后做JS层面的循环分组,数据量超过1万条就很容易触发Function内存超限、执行超时,上述原生聚合逻辑会将计算下推到Ontology存储层,百万级数据量也能稳定返回 - 如果你的Ontology中属性API名做了驼峰/下划线转换,比如
hours worked被自动转成hoursWorked,直接把代码里的a["hours worked"]替换成a.hoursWorked即可,属性名可以在Ontology对象类型的属性详情页查看对应的API标识 - 如果需要过滤掉工时为空、部门/班次未填写的异常数据,可以在filter阶段追加对应的非空判断逻辑,比如
.filter(item => item.department.notNull())
内容的提问来源于stack exchange,提问作者kevpl541991
相关产品推荐
相关产品推荐

