如何通过REST API为Palantir Foundry数据集上传的CSV应用Schema
问题:通过Foundry REST API上传CSV后无法自动应用Schema
我用Foundry REST API上传CSV文件到数据集时,只能完成文件上传,但没法把文件内容填充到实际数据集里。网页端可以通过「Apply a schema」操作实现这个效果,但我需要用Python做程序化处理。
我现有的文件上传函数能成功上传文件,但缺少Schema应用步骤:
def upload_file_to_foundry( input_file, token, dataset_rid, base_url): headers = { "content-type": "application/octet-stream", "authorization": "Bearer {}".format(token) } requests.post(f'{base_url}/api/v1/datasets/{dataset_rid}/files:upload?filePath=myUploadFile.csv', headers=headers, data=open(input_file, 'rb'))
同事给了一段旧代码,原本是用来上传CSV并自动应用Schema的,但调用Foundry Schema推理接口时返回404错误:
def query_foundry(url, token, headers=None, #attrHeaders=None, params=None, data=None, attrJson=None, files=None, stream=False ) -> (requests.Response): ''' purpose: general post request used by various foundry queries parameters: url - url used for the post request token - bearer token from foundry (auth token) headers - set the auth and content-type for the post attrHeaders - for providing additional headers params - send params as a query string, similar to a get request data - send payload as the body of the request attrJson - send payload as the body of the request, taken in form of json files - send payload as part of the body with a content-type of multipart/form-data stream - stream the request (default: False) ''' # prepare and execute POST request if not headers: headers = {} headers["Authorization"] = f'Bearer {token}' headers["Content-Type"] = 'application/json' response = requests.post(url, headers=headers, params=params, json=attrJson, data=data, files=files, stream=stream ) return response def csv_to_foundry_dataset(token, datasetRID, transactionRID, fileNameUpload, branch='master' ): ''' purpose: upload csv to foundry, commit the transaction and set schema based on the foundry-schema-inference api parameters: token - bearer token from foundry (auth token) datasetRID - the rid of the dataset to overwrite transactionRID - the rid of the active transaction for the dataset fileNameUpload - location of the csv to upload branch - the branch of the target dataset (defualt: master) ''' headers = {} headers["Authorization"] = f'Bearer {token}' # strip the byte order mask from the file and encode as utf-8. this is necessary # when downloading a csv from foundry as it adds a BOM to the beginning of file bom_file = open(fileNameUpload, mode='r', encoding='utf-8-sig').read() open(fileNameUpload, mode='w', encoding='utf-8').write(bom_file) url = f"{baseUrl}/foundry-data-proxy/api/dataproxy/datasets/{datasetRID}/transactions/{transactionRID}" files = {'upload': ('data.csv', open(fileNameUpload, 'r'), 'csv')} # upload the csv to foundry response = query_foundry(url=url, token=token, headers=headers, files=files) if response.ok: # commit the transaction url = f"{baseUrl}/foundry-catalog/api/catalog/datasets/{datasetRID}/transactions/{transactionRID}/commit" response = query_foundry(url=url, token=token, data="{}") # 此处返回404的代码段 # get infered schema from foundry url = f"{baseUrl}/foundry-schema-inference/api/datasets/{datasetRID}/branches/{branch}/schema" response = query_foundry(url=url, token=token, data="{}") c = json.loads(response.content) schema = c["data"]["foundrySchema"] # apply schema to foundry dataset url = f"{baseUrl}/foundry-metadata/api/schemas/datasets/{datasetRID}/branches/{branch}" response = query_foundry(url=url, token=token, attrJson=schema) else: response.raise_for_status() return response
出现404的原因是旧代码调用的foundry-schema-inference接口已经被废弃或路径变更,以下是修复后的完整解决方案:
修复方案:使用最新Foundry API流程
要完成CSV上传+Schema应用,需要遵循数据集事务操作的完整流程,同时使用当前有效的Schema推理接口:
步骤说明
- 创建数据集事务(替代直接上传文件)
- 在事务中上传CSV文件
- 提交事务
- 使用Foundry官方的Schema推理接口获取自动生成的Schema
- 将推理得到的Schema应用到数据集分支
完整代码实现
import requests import json def upload_csv_with_schema(token, dataset_rid, input_file, base_url, branch="master"): # 1. 创建数据集事务 create_transaction_url = f"{base_url}/api/v1/datasets/{dataset_rid}/transactions" headers = { "Authorization": f"Bearer {token}", "Content-Type": "application/json" } transaction_response = requests.post( create_transaction_url, headers=headers, json={"branchId": branch, "operation": "OVERWRITE"} ) transaction_response.raise_for_status() transaction_rid = transaction_response.json()["rid"] # 2. 在事务中上传CSV文件 upload_url = f"{base_url}/api/v1/datasets/{dataset_rid}/transactions/{transaction_rid}/files:upload?filePath=data.csv" with open(input_file, "rb") as f: upload_response = requests.post( upload_url, headers={"Authorization": f"Bearer {token}", "Content-Type": "application/octet-stream"}, data=f ) upload_response.raise_for_status() # 3. 提交事务 commit_url = f"{base_url}/api/v1/datasets/{dataset_rid}/transactions/{transaction_rid}:commit" commit_response = requests.post(commit_url, headers=headers) commit_response.raise_for_status() # 4. 自动推理Schema(使用当前有效接口) infer_schema_url = f"{base_url}/api/v1/datasets/{dataset_rid}/branches/{branch}/schema:infer" infer_response = requests.post( infer_schema_url, headers=headers, json={"sourceType": "FILE", "filePath": "data.csv"} ) infer_response.raise_for_status() inferred_schema = infer_response.json()["schema"] # 5. 应用Schema到数据集 apply_schema_url = f"{base_url}/api/v1/datasets/{dataset_rid}/branches/{branch}/schema" apply_response = requests.put( apply_schema_url, headers=headers, json=inferred_schema ) apply_response.raise_for_status() return apply_response # 使用示例 # upload_csv_with_schema( # token="your_foundry_token", # dataset_rid="ri.foundry.main.dataset.xxxxxx", # input_file="your_local_file.csv", # base_url="https://your-foundry-instance.com" # )
关键修复点
- 替换了旧的废弃接口,使用Foundry V1 API官方推荐的
/schema:infer端点 - 遵循标准的「事务创建→文件上传→事务提交」流程,替代旧的代理接口调用
- 统一使用官方V1 API路径,避免因接口版本迭代导致的404问题
注意事项
- 确保你的Foundry API令牌拥有数据集编辑和Schema管理权限
- 如果CSV带有BOM头,可以保留旧代码中的BOM移除逻辑
- 若需要自定义Schema,可跳过自动推理步骤,直接构造符合Foundry Schema格式的JSON对象并应用
内容的提问来源于stack exchange,提问作者baobobs
相关产品推荐
相关产品推荐

