You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Vision OCR技术问询:inputConfig是否支持本地存储等非Cloud Storage数据源

Google Vision OCR: Can You Skip Cloud Storage for Local Files?

Hey there! Let me break this down clearly for you: you can’t modify the inputConfig to directly support local file storage for PDF OCR with Google Vision. Here’s why, plus practical workarounds to handle local PDFs smoothly:

Why Local Files Aren’t Supported for PDFs

Google Vision’s batch document OCR (built for multi-page files like PDFs) requires input files to be hosted on Google Cloud Storage. The gcsSource field in inputConfig is mandatory for PDF processing—this is because the Vision service needs persistent, server-side access to process all pages efficiently, and it can’t reach files stored on your local machine or external storage systems directly.

Workarounds to Handle Local PDFs

Even though you can’t skip Cloud Storage entirely, you can automate the process to make it feel seamless:

1. Use Byte Streams for Single-Page Images (Not PDFs)

If you’re working with single-page image files (JPG, PNG, etc.), you can send the file’s raw bytes directly in the API request using the content field instead of gcsSource. This skips Cloud Storage entirely, but note this method doesn’t support multi-page PDFs.

2. Automate Upload + OCR + Cleanup with SDKs

For PDFs, the most straightforward approach is to use Google Cloud’s client SDKs to handle temporary uploads to Cloud Storage, run the OCR, and delete the file afterward if you don’t need it. Here’s a quick Python example to demonstrate the workflow:

from google.cloud import vision
from google.cloud import storage

# Initialize clients
storage_client = storage.Client()
vision_client = vision.ImageAnnotatorClient()

# Configuration
bucket_name = "your-cloud-storage-bucket"
local_pdf_path = "/path/to/your/local/file.pdf"
temp_blob_name = "temp-processing.pdf"
output_gcs_uri = f"gs://{bucket_name}/ocr-results/"

# Step 1: Upload local PDF to Cloud Storage
bucket = storage_client.bucket(bucket_name)
temp_blob = bucket.blob(temp_blob_name)
temp_blob.upload_from_filename(local_pdf_path)

# Step 2: Set up OCR request
input_config = {
    "gcs_source": {"uri": f"gs://{bucket_name}/{temp_blob_name}"},
    "mime_type": "application/pdf"
}
features = [{"type_": vision.Feature.Type.DOCUMENT_TEXT_DETECTION}]
output_config = {"gcs_destination": {"uri": output_gcs_uri}}

# Step 3: Run asynchronous OCR
async_request = vision.AsyncAnnotateFileRequest(
    requests=[{
        "input_config": input_config,
        "features": features,
        "output_config": output_config
    }]
)
operation = vision_client.async_batch_annotate_files(requests=[async_request])
print("Waiting for OCR to complete...")
operation.result(timeout=180)  # Wait for processing to finish

# Step 4: Optional - Clean up the temporary PDF file
temp_blob.delete()
print("Temporary file deleted, OCR complete!")

This script handles the entire workflow automatically—no manual uploads required. Cloud Storage’s temporary storage costs are minimal, and deleting the file right after processing keeps your storage tidy.

内容的提问来源于stack exchange,提问作者Deni Mardiana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:16:19