You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Excel上传器开发:如何为每条数据添加文档ID?

Hey there! Let's walk through how to add this document ID association to your Excel uploader. Here's a clear, actionable approach with code examples to make it concrete:

Step 1: Adjust Your Workflow Order

First, you’ll need to reorder your existing process to prioritize document upload before data processing:

  • Original flow: Extract data from Excel → Send data to API
  • Updated flow: Upload Excel document to API → Retrieve document ID from response → Attach ID to every data record → Send records to API
Step 2: Upload the Excel Document & Fetch Its ID

Start by uploading your Excel file to the document storage API, then extract the unique ID from the successful response. Below’s a Python example using requests (adapt this to your tech stack if needed):

import requests

# Define your API endpoints and file path
DOC_UPLOAD_API = "https://your-api-endpoint.com/documents"
EXCEL_FILE_PATH = "./your_uploaded_file.xlsx"

# Upload the Excel file
with open(EXCEL_FILE_PATH, "rb") as file:
    upload_payload = {
        "file": ("data.xlsx", file, "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet")
    }
    response = requests.post(DOC_UPLOAD_API, files=upload_payload)

# Handle response and extract document ID
if response.status_code in (200, 201):
    # Replace "document_id" with the actual key from your API's response
    document_id = response.json()["document_id"]
    print(f"Document uploaded successfully! ID: {document_id}")
else:
    # Add error handling (logging, retries, user alerts) here
    raise Exception(f"Failed to upload document: {response.text}")
Step 3: Attach Document ID to Each Data Record

Next, read your Excel data as you already do, then add the document ID as a new field to every record. Using pandas for Excel parsing (again, adjust to your tooling):

import pandas as pd

# Read Excel data into a DataFrame
df = pd.read_excel(EXCEL_FILE_PATH)

# Add the document ID as a new column to all rows
df["document_id"] = document_id

# Convert DataFrame to a list of dictionaries (ready for API submission)
data_records = df.to_dict("records")
Step 4: Send Records with Document ID to the Data API

Finally, send the enriched data records to your storage API just like you did before—now each record includes the parent document’s ID:

DATA_STORAGE_API = "https://your-api-endpoint.com/data-records"
headers = {"Content-Type": "application/json"}

response = requests.post(DATA_STORAGE_API, json=data_records, headers=headers)

if response.status_code == 200:
    print("All records saved with document ID successfully!")
else:
    raise Exception(f"Failed to save data records: {response.text}")
Key Considerations for Production
  • Error Handling: Add robust error handling for failed uploads (e.g., retry logic for transient API errors, logging for debugging)
  • Large Files: If dealing with large Excel files, process data in chunks instead of loading everything into memory to avoid performance issues
  • API Contract Alignment: Double-check that the document ID’s data type (string/integer) matches what your data API expects
  • Idempotency: If retries are needed, ensure your document upload API supports idempotent requests to avoid duplicate document entries

内容的提问来源于stack exchange,提问作者Chuck Villavicencio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:10:19