You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言paws包Textract能否不使用S3Object直接读取本地PDF文件

解答

不可以,AWS Textract 的 start_document_analysis 异步接口本身不支持直接传入本地文件路径,仅支持读取存储在 AWS S3 桶中的文档。

所有 Textract 异步处理类接口(包括 start_document_analysis、start_document_text_detection 等)的设计逻辑就是仅接受 S3 存储的文件作为输入,DocumentLocation 参数仅支持填写 S3 对象信息,没有本地文件路径的传参入口。

如果需要处理本地文件,可以根据使用场景选择对应方案:

  • 处理单页文档(单页PDF、JPG/PNG等格式图片):改用同步接口 analyze_document,该接口支持直接传入本地文件的字节流,示例代码如下:
# 读取本地文件为raw格式
local_file_path <- "本地文件路径.pdf"
file_content <- readBin(local_file_path, "raw", n = file.info(local_file_path)$size)

# 调用同步分析接口
analysis_result <- textract$analyze_document(
  Document = list(Bytes = file_content),
  FeatureTypes = c("TABLES", "FORMS") # 按需选择需要提取的内容类型
)
  • 处理多页PDF必须使用异步接口:先将本地PDF上传到S3桶,再调用 start_document_analysis 传入对应S3路径,示例代码如下:
# 初始化S3客户端
s3_client <- paws::s3()
bucket_name <- "你的S3桶名"
s3_file_key <- "文件在S3中存储的名称.pdf"

# 上传本地文件到S3
s3_client$put_object(
  Bucket = bucket_name,
  Key = s3_file_key,
  Body = "本地PDF文件路径.pdf"
)

# 调用异步分析接口
textract$start_document_analysis(
  DocumentLocation = list(
    S3Object = list(Bucket = bucket_name, Name = s3_file_key)
  ),
  FeatureTypes = c("TABLES", "FORMS")
)

内容的提问来源于stack exchange,提问作者jkortner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 15:45:01