Palantir Foundry离线运行Spark-NLP库遇文件未找到问题求助
在Palantir Foundry中离线运行Spark-NLP的问题
问题背景
我正尝试在Palantir Foundry中离线运行Spark-NLP库,由于环境未配置网络出口,无法发起HTTP请求,因此通过Spark-NLP模型中心下载explain_document_dl模型包,以离线模式使用该库。这只是Spark-NLP快速入门中的简单示例,目的仅为使其在Foundry中正常运行。
环境说明
- 使用PYPI上的Spark-NLP 5.1.2版本
- 已按照Palantir公开文档中「同时需要conda包和jar包的库」的特殊说明加载Spark-NLP库,确认该场景具备可行性
尝试过的操作
- 此前通过从Foundry数据集文件系统加载/解压zip包的方式成功运行BERTopic,因此尝试将该方法应用到Spark-NLP
- Spark-NLP 5.1版本的模型中心中有6个不同版本的explain_document_dl模型,除版本号外无法明确差异,尝试加载所有版本均未成功
- 尝试执行
pipeline = PretrainedPipeline('explain_document_dl', lang='en')加载预训练模型,试图查看其HTTP请求获取的具体版本,但未成功 - 不确定问题出在Foundry、Spark-NLP离线运行方式,还是两者皆有
代码实现
from transforms.api import transform, Input, Output from sparknlp.base import PipelineModel # from sparknlp.annotator import * from sparknlp.pretrained import PretrainedPipeline # import sparknlp from zipfile import ZipFile import tempfile import os import shutil import base64 def download_file(filesystem, input_file_path, local_file_path=None, base64_decode=False): """ Download a file from a Foundry dataset to the local filesystem. If the input_file_path is None, a temporary file is created, which you must delete yourself after using it. :param filesystem: an instance of transform.api.FileSystem :param input_file_path: logical path on the Foundry dataset to download from :param local_file_path: path of the file to download to on the local file system (default=None) :base64_encode: if set to True, decode data using base64 (default=False) :return: str path of the downloaded file on the local file system """ # Check if a different temp directory is specified in the Spark environment, and use it if so TEMPDIR_OVERRIDE = os.getenv('FOUNDRY_PYTHON_TEMPDIR') tempfile.tempdir = TEMPDIR_OVERRIDE if TEMPDIR_OVERRIDE is not None else tempfile.tempdir if local_file_path is None: _, local_file_path = tempfile.mkstemp() if base64_decode: _, download_file_path = tempfile.mkstemp() else: download_file_path = local_file_path try: with filesystem.open(input_file_path, 'rb') as f_in, open(download_file_path, 'wb') as f_out: shutil.copyfileobj(f_in, f_out) if base64_decode: with open(download_file_path, 'rb') as fin, open(local_file_path, 'wb') as fout: base64.decode(fin, fout) return local_file_path finally: if base64_decode: os.remove(download_file_path) @transform( raw_model=Input("ri.foundry.main.dataset.6b9128f7-dc4b-4c99-81c8-94fb4b0f9ab4"), model_output=Output("ri.foundry.main.dataset.67b42a35-5105-41ef-83f4-b9ccfd347d93"), ) def compute(raw_model, model_output): # Offline mode temp_dir = tempfile.mkdtemp() model_package = download_file( raw_model.filesystem(), "explain_document_dl_en_4.4.2_3.2_1685186531034.zip", "{}/explain_document_dl_en_4.4.2_3.2_1685186531034.zip".format(temp_dir) ) with ZipFile(model_package, "r") as zObject: # Extract all files from zip and put them in temp directory explain_document_dl = tempfile.mkdtemp() zObject.extractall(path=explain_document_dl) pipeline = PipelineModel.load(explain_document_dl) # Your testing dataset text = "The Mona Lisa is a 16th century oil painting created by Leonardo. It's held at the Louvre in Paris." # Annotate your testing dataset result = pipeline.annotate(text) # What's in the pipeline list(result.keys()) # Check the results result['entities'] model_output.write_dataframe(result) zObject.close()
错误信息
{ "errorCode": "CUSTOM_CLIENT", "errorName": "Spark:JobAborted", "errorInstanceId": "d0ec47f9-628c-411c-be10-ac3b1c81b941", "safeArgs": { "pythonVersion": "3.10.12", "exceptionClass": "java.io.FileNotFoundException", "message": "ServiceException: CUSTOM_CLIENT (Spark:JobAborted)" }, "unsafeArgs": { "stacktrace": "org.apache.spark.SparkException: Job aborted due to stage failure: Task 0 in stage 1.0 failed 4 times, most recent failure: Lost task 0.3 in stage 1.0 (TID 4) : java.io.FileNotFoundException: File file:/tmp/tmp67p250jr/metadata/part-00000 does not exist\n\tat" }
内容的提问来源于stack exchange,提问作者tessa
相关产品推荐
相关产品推荐

