You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python读取SharePoint Online(Office365)中的PDF和Word文件

如何用Python读取SharePoint Online(Office365)中的PDF和Word文件?

要读取SharePoint Online中的PDF和Word文件,核心是先通过API获取文件字节流,再复用你提供的本地文件解析逻辑处理流数据。以下是完整实现方案:

1. 安装依赖库

先安装所需工具包:

pip install office365-rest-python-client PyPDF2 python-docx

2. 完整代码实现

核心逻辑:

  • 用office365-rest-python-client建立SharePoint站点连接
  • 下载目标文件的字节流
  • 调用解析函数提取内容(复用你提供的本地读取逻辑)
import io
from office365.sharepoint.client_context import ClientContext
from office365.runtime.auth.user_credential import UserCredential
from PyPDF2 import PdfReader
from docx import Document

def extract_file_content(file_bytes, extension):
    """从字节流提取文件内容,复用本地解析逻辑"""
    file_contents = ""
    if extension == "pdf":
        file_stream = io.BytesIO(file_bytes)
        pdf_reader = PdfReader(file_stream)
        for page in pdf_reader.pages:
            extracted_text = page.extract_text()
            if extracted_text:
                file_contents += extracted_text
    elif extension == "docx":
        file_stream = io.BytesIO(file_bytes)
        doc = Document(file_stream)
        paragraphs = [p.text for p in doc.paragraphs if p.text]
        file_contents = "\n".join(paragraphs)
    return file_contents

def get_sharepoint_file_content(site_url, username, password, relative_file_path):
    """从SharePoint Online获取文件并提取内容"""
    # 初始化SharePoint连接
    ctx = ClientContext(site_url).with_credentials(UserCredential(username, password))
    # 获取目标文件对象
    file = ctx.web.get_file_by_server_relative_path(relative_file_path)
    # 下载文件字节流
    file_bytes = file.read().execute_query()
    # 解析文件扩展名
    extension = relative_file_path.split(".")[-1].lower()
    # 提取文件内容
    return extract_file_content(file_bytes, extension)

# 示例调用(替换为你的实际配置)
if __name__ == "__main__":
    SITE_URL = "https://你的租户.sharepoint.com/sites/你的站点"
    USERNAME = "你的账号@你的租户.onmicrosoft.com"
    PASSWORD = "你的密码或应用专用密码"
    # 文件相对路径(从站点根目录开始,比如"/Shared Documents/测试文档.pdf")
    FILE_PATH = "/Shared Documents/测试报告.docx"

    content = get_sharepoint_file_content(SITE_URL, USERNAME, PASSWORD, FILE_PATH)
    print(content)

注意事项

  • 若账号开启MFA(多因素认证),不能直接用密码登录,需改用应用权限或证书认证,可参考office365-rest-python-client官方配置说明
  • 确保账号对目标文件拥有读取权限
  • 大文件建议分块下载,上述代码适用于中小文件场景

内容的提问来源于stack exchange,提问作者PramoD19

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 00:22:16