如何通过Google Drive API上传时禁用OCR功能
Good news—yes, the Google Drive v3 API does support disabling OCR, and the fix for your issue is actually pretty straightforward once you spot the root cause in your code.
The Root of Your OCR Problem
Looking at your MIMETYPE_MAPPING dictionary, you've mapped all image types (like image/png, image/jpeg, etc.) to application/vnd.google-apps.document. When you upload an image with this MIME type, Google Drive automatically tries to convert the image into an editable Google Doc—and that conversion process includes default OCR to extract text from the image. That's exactly why you're seeing unwanted OCR runs.
Two Ways to Disable OCR
1. Stop Mapping Images to Google Docs (Recommended)
If your goal is just to convert images to PDF without extracting editable text, you don't need to turn images into Google Docs at all. Add a check in your drive_upload function to keep original image MIME types, skipping the Google Doc conversion that triggers OCR:
def drive_upload(fp, fn): """ Uploads the file found at fp to root of google drive account with name fn Returns the id of the new file """ mimetype = magic.from_file(fp, mime=True) drive_service = get_service() # Keep original MIME type for images to avoid OCR if mimetype.startswith('image/'): file_metadata = { 'name': fn, 'mimeType': mimetype } else: file_metadata = { 'name': fn, 'mimeType': MIMETYPE_MAPPING.get(mimetype, 'application/vnd.google-apps.document'), } media = MediaFileUpload(fp, mimetype=mimetype, resumable=True) file = drive_service.files().create(body=file_metadata, media_body=media, fields='id').execute() return file.get('id')
Drive still lets you export images directly to PDF via files.export_media, so this won't break your conversion workflow—you just skip the unnecessary (and OCR-heavy) Google Doc conversion step.
2. Force Disable OCR (If You Need Google Doc Conversion)
If you do need to convert images to Google Docs but want to skip OCR, you can use the contentHints field in your file metadata to override Drive's default behavior. By setting an empty indexableText, you tell Drive not to run OCR (it'll use your provided empty text instead):
file_metadata = { 'name': fn, 'mimeType': MIMETYPE_MAPPING.get(mimetype, 'application/vnd.google-apps.document'), 'contentHints': { 'indexableText': '', } }
This creates a Google Doc linked to your image but without extracting any text via OCR.
Either approach will stop Drive from automatically running OCR on your uploaded images.
内容的提问来源于stack exchange,提问作者frederick_otus

