You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

通过Python上传至S3的PDF无法在浏览器预览,需设置哪些参数?

问题

我已经能通过boto3客户端以编程方式对AWS S3存储桶中的PDF文件执行高亮、上传及下载操作,但希望实现PDF文件在浏览器中的直接预览。通过AWS控制台上传文件时,其公开URL支持直接预览;但通过AWS SDK编程上传后,使用公开URL会触发文件下载而非预览。请问在编程上传时需设置哪些参数,才能让PDF在浏览器中正常预览?

以下是我的代码片段:

# Accessing files from s3 my_s3_bucket
import boto3
from PyPDF2 import PdfFileReader
from io import BytesIO
import os
import fitz


aws_access_key_id = os.environ['aws_access_key_id']
aws_secret_access_key = os.environ['aws_secret_access_key']
aws_region_name = os.environ['aws_region_name']

session = boto3.Session(profile_name='default')
s3 = session.client('s3')

print('Boto client created successfully')

bucket_name = 'my-bucket'
item_name = 'Sample.pdf'

s3_object = s3.get_object(Bucket=bucket_name,Key=item_name)
body = s3_object['Body']

fs = body.read()
pdf=fitz.open("pdf", stream=BytesIO(fs))
pdfPage = pdf[42]
    
es_highlighted_result = ['line 1', 
                         'line 2' , 
                         'line 3']
for line in es_highlighted_result:
    r1 = pdfPage.search_for(line)
    pdfPage.addHighlightAnnot(r1)

output_buffer = BytesIO()
pdf.save(output_buffer)
output_filepath = 'my-path-to-file'
output_file = output_filepath + item_name
new_bytes = pdf.write()

#----------------------------------------
# Logic to write the file on local and then upload
# ------------------------------------
with open(output_file,mode='wb') as f:
        f.write(output_buffer.getbuffer())

upload_file_bucket = bucket_name
upload_file_key = str(output_file)
s3.upload_file(output_file,upload_file_bucket,upload_file_key)

#------Tried this also --------------
# s3.Bucket(bucket_name).put_object(Key=output_file, Body=new_bytes)
# s3.put_object(Bucket=bucket_name,Key=output_file,Body=new_bytes,ACL='public-read')
#----------------------------------------
print('File uploaded successfully')

url = s3.generate_presigned_url(ClientMethod='get_object', 
    Params={'Bucket': bucket_name, 'Key': output_file},ExpiresIn=3600)

print(url)
解决方案

问题出在文件的Content-Type元数据上。AWS控制台上传时会自动识别PDF文件,设置Content-Type: application/pdf;但SDK编程上传默认不会自动配置这个元数据,导致浏览器无法识别文件类型,只能触发下载。

你需要在上传时显式指定Content-Type参数,同时确保文件访问权限符合需求(如果用公开URL的话)。以下是两种上传方式的修改方案:

1. 修改upload_file方法调用

调用upload_file时添加ExtraArgs参数,指定Content-Type和访问权限:

s3.upload_file(
    output_file,
    upload_file_bucket,
    upload_file_key,
    ExtraArgs={
        'ContentType': 'application/pdf',
        'ACL': 'public-read'  # 若需要公开访问则添加
    }
)

2. 修改put_object方法调用

如果直接用字节流上传,同样需要指定Content-Type:

s3.put_object(
    Bucket=bucket_name,
    Key=output_file,
    Body=new_bytes,
    ACL='public-read',
    ContentType='application/pdf'
)

额外优化:预签名URL强制预览

如果你使用预签名URL访问文件,还可以在生成URL时添加ResponseContentDisposition参数,强制浏览器以内联方式显示文件,作为双重保障:

url = s3.generate_presigned_url(
    ClientMethod='get_object',
    Params={
        'Bucket': bucket_name,
        'Key': output_file,
        'ResponseContentDisposition': 'inline'
    },
    ExpiresIn=3600
)

内容的提问来源于stack exchange,提问作者Dinesh Zende

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 05:55:46