You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Palantir Foundry离线运行Spark-NLP库遇文件未找到问题求助

在Palantir Foundry中离线运行Spark-NLP的问题

问题背景

我正尝试在Palantir Foundry中离线运行Spark-NLP库,由于环境未配置网络出口,无法发起HTTP请求,因此通过Spark-NLP模型中心下载explain_document_dl模型包,以离线模式使用该库。这只是Spark-NLP快速入门中的简单示例,目的仅为使其在Foundry中正常运行。

环境说明

  • 使用PYPI上的Spark-NLP 5.1.2版本
  • 已按照Palantir公开文档中「同时需要conda包和jar包的库」的特殊说明加载Spark-NLP库,确认该场景具备可行性

尝试过的操作

  • 此前通过从Foundry数据集文件系统加载/解压zip包的方式成功运行BERTopic,因此尝试将该方法应用到Spark-NLP
  • Spark-NLP 5.1版本的模型中心中有6个不同版本的explain_document_dl模型,除版本号外无法明确差异,尝试加载所有版本均未成功
  • 尝试执行pipeline = PretrainedPipeline('explain_document_dl', lang='en')加载预训练模型,试图查看其HTTP请求获取的具体版本,但未成功
  • 不确定问题出在Foundry、Spark-NLP离线运行方式,还是两者皆有

代码实现

from transforms.api import transform, Input, Output
from sparknlp.base import PipelineModel
# from sparknlp.annotator import *
from sparknlp.pretrained import PretrainedPipeline
# import sparknlp
from zipfile import ZipFile
import tempfile
import os
import shutil
import base64


def download_file(filesystem, input_file_path, local_file_path=None, base64_decode=False):
    """
    Download a file from a Foundry dataset to the local filesystem.
    If the input_file_path is None, a temporary file is created, which you must delete yourself after using it.
    :param filesystem: an instance of transform.api.FileSystem
    :param input_file_path: logical path on the Foundry dataset to download from
    :param local_file_path: path of the file to download to on the local file system (default=None)
    :base64_encode: if set to True, decode data using base64 (default=False)
    :return: str path of the downloaded file on the local file system
    """

    # Check if a different temp directory is specified in the Spark environment, and use it if so
    TEMPDIR_OVERRIDE = os.getenv('FOUNDRY_PYTHON_TEMPDIR')
    tempfile.tempdir = TEMPDIR_OVERRIDE if TEMPDIR_OVERRIDE is not None else tempfile.tempdir

    if local_file_path is None:
        _, local_file_path = tempfile.mkstemp()

    if base64_decode:
        _, download_file_path = tempfile.mkstemp()
    else:
        download_file_path = local_file_path

    try:
        with filesystem.open(input_file_path, 'rb') as f_in, open(download_file_path, 'wb') as f_out:
            shutil.copyfileobj(f_in, f_out)

        if base64_decode:
            with open(download_file_path, 'rb') as fin, open(local_file_path, 'wb') as fout:
                base64.decode(fin, fout)

        return local_file_path

    finally:
        if base64_decode:
            os.remove(download_file_path)


@transform(
    raw_model=Input("ri.foundry.main.dataset.6b9128f7-dc4b-4c99-81c8-94fb4b0f9ab4"),
    model_output=Output("ri.foundry.main.dataset.67b42a35-5105-41ef-83f4-b9ccfd347d93"),
)
def compute(raw_model, model_output):
    # Offline mode
    temp_dir = tempfile.mkdtemp()
    model_package = download_file(
        raw_model.filesystem(),
        "explain_document_dl_en_4.4.2_3.2_1685186531034.zip",
        "{}/explain_document_dl_en_4.4.2_3.2_1685186531034.zip".format(temp_dir)
    )
    with ZipFile(model_package, "r") as zObject:
        # Extract all files from zip and put them in temp directory
        explain_document_dl = tempfile.mkdtemp()
        zObject.extractall(path=explain_document_dl)

        pipeline = PipelineModel.load(explain_document_dl)

    # Your testing dataset
    text = "The Mona Lisa is a 16th century oil painting created by Leonardo. It's held at the Louvre in Paris."

    # Annotate your testing dataset
    result = pipeline.annotate(text)

    # What's in the pipeline
    list(result.keys())

    # Check the results
    result['entities']
    model_output.write_dataframe(result)
    zObject.close()

错误信息

{
"errorCode": "CUSTOM_CLIENT",
"errorName": "Spark:JobAborted",
"errorInstanceId": "d0ec47f9-628c-411c-be10-ac3b1c81b941",
"safeArgs": {
"pythonVersion": "3.10.12",
"exceptionClass": "java.io.FileNotFoundException",
"message": "ServiceException: CUSTOM_CLIENT (Spark:JobAborted)"
},
"unsafeArgs": {
"stacktrace": "org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 1.0 failed 4 times, most recent failure: Lost task 0.3 in stage 1.0 (TID 4) : java.io.FileNotFoundException: File file:/tmp/tmp67p250jr/metadata/part-00000 does not exist\n\tat"
}

内容的提问来源于stack exchange,提问作者tessa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 22:42:09