如何在Pentaho用户定义Java类中彻底删除指定输入字段
在PDI用户自定义Java类中彻底删除指定字段的解决方案
看起来你已经搞定了字段值的拼接,但卡在了彻底删除字段这一步——你当前的代码只是把要删除的字段设为空字符串,但PDI的行结构是由RowMeta(行元数据)和RowData(行数据)共同决定的,只清空数据不会从元数据里移除字段,所以它依然会出现在输出中。
下面我会帮你修改代码,实现真正删除字段B,同时保留A和更新后的C:
核心思路
要彻底删除字段,需要同步完成两件事:
- 构建新的输出行元数据(RowMeta),只保留你需要的字段(A和C)
- 构建对应的输出行数据(RowData),和新的元数据字段一一对应
修改后的完整代码
private String outFieldName1; // 对应字段A private String outFieldName2; // 对应字段C private String removeFieldName; // 对应字段B private int outFieldIndex1; private int outFieldIndex2; private int removeFieldIndex; // 新增:存储自定义的输出行元数据 private RowMetaInterface outputRowMeta; public boolean processRow(StepMetaInterface smi, StepDataInterface sdi) throws KettleException { Object[] inputRow = getRow(); if (inputRow == null) { setOutputDone(); return false; } if (first) { processMetadata(sdi); first = false; } // 拼接B和C的值,准备写入C字段 String mergedContent = get(Fields.In, removeFieldName).getString(inputRow) + " " + get(Fields.In, outFieldName2).getString(inputRow); pushOutputRow(inputRow, mergedContent); return true; } private void processMetadata(StepDataInterface sdi) throws KettleException { outFieldName1 = getParameter("OUT1"); outFieldName2 = getParameter("OUT2"); removeFieldName = getParameter("REMOVE"); RowMetaInterface inputRowMeta = getInputRowMeta(); outFieldIndex1 = inputRowMeta.indexOfValue(outFieldName1); outFieldIndex2 = inputRowMeta.indexOfValue(outFieldName2); removeFieldIndex = inputRowMeta.indexOfValue(removeFieldName); // 构建输出行元数据:复制输入元数据后移除目标字段 outputRowMeta = (RowMetaInterface) inputRowMeta.clone(); outputRowMeta.removeValue(removeFieldIndex); // 将新元数据绑定到StepData,确保后续输出使用正确的元数据 ((BaseStepData)sdi).outputRowMeta = outputRowMeta; } private void pushOutputRow(Object[] inputRow, String mergedContent) throws KettleException { // 输出行长度严格匹配新元数据的字段数量 Object[] outRow = RowDataUtil.allocateRowData(outputRowMeta.size()); int outIndex = 0; for (int inIndex = 0; inIndex < inputRow.length; inIndex++) { // 跳过要删除的字段B,不写入输出行 if (inIndex == removeFieldIndex) { continue; } // 对字段C赋值拼接后的内容,其他字段直接复制原数据 if (inIndex == outFieldIndex2) { outRow[outIndex] = mergedContent; } else { outRow[outIndex] = inputRow[inIndex]; } outIndex++; } // 使用新的元数据输出行数据 putRow(outputRowMeta, outRow); }
关键修改点说明
重构输出元数据:
- 用
inputRowMeta.clone()复制输入元数据,避免修改原数据影响其他步骤 - 调用
removeValue(removeFieldIndex)彻底移除字段B的元信息,这样输出字段列表里就只剩A和C了 - 将新元数据绑定到StepData,确保后续
putRow操作使用正确的元数据模板
- 用
同步输出行数据:
- 不再按输入行的索引直接赋值,而是遍历输入字段,跳过要删除的字段,逐个填充到输出行
- 对需要更新的字段C,替换为拼接后的内容;其他字段(如A)直接复制原数值
避免数据与元数据不匹配:
- 输出行的长度严格对应新元数据的字段数,彻底解决了原代码中“字段存在但值为空”的问题
测试验证
运行转换后,你可以通过PDI的预览功能查看输出字段列表,确认字段B已经被完全移除,而非仅被清空。如果是更复杂的多字段删除场景,只需要在processMetadata中批量移除对应字段即可。
内容的提问来源于stack exchange,提问作者Mikhail Ovchinnikov
相关产品推荐
相关产品推荐

