You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python结合AAD应用访问SharePoint文件及子文件夹?

实现Python通过AAD应用访问SharePoint文件及子文件夹(无需本地下载)

完全可以实现这类需求,核心是通过Microsoft Graph API或SharePoint REST API结合AAD应用权限认证,直接在内存中读取、操作文件内容,无需下载到本地。以下是具体方案和工具:

核心实现逻辑

  1. 在Azure门户注册AAD应用,获取租户ID、客户端ID、客户端密钥(或证书)。
  2. 为AAD应用配置SharePoint相关的应用权限(如Sites.Read.All、Files.Read.All),并获取管理员同意。
  3. 通过认证获取访问令牌,调用API遍历文件夹、读取文件内容,全程在内存中处理数据。

推荐Python包

1. office365-rest-python-client

专门针对Office 365生态的封装库,简化了SharePoint的API调用,无需手动处理请求细节,适合快速开发。

2. azure-identity + requests

azure-identity负责AAD认证令牌的获取,requests直接调用Microsoft Graph API,灵活性更高,适合自定义复杂操作场景。

代码示例

使用office365-rest-python-client

from office365.sharepoint.client_context import ClientContext
from office365.runtime.auth.client_credential import ClientCredential

# 配置参数
tenant_id = "你的租户ID"
client_id = "你的AAD应用客户端ID"
client_secret = "你的AAD应用客户端密钥"
site_url = "https://你的租户.sharepoint.com/sites/目标站点名称"

# 初始化认证上下文
ctx = ClientContext(site_url).with_credentials(ClientCredential(client_id, client_secret))

# 递归遍历文件夹并读取文件内容
def traverse_folder(folder_rel_url):
    folder = ctx.web.get_folder_by_server_relative_url(folder_rel_url)
    files = folder.files
    sub_folders = folder.folders
    ctx.load(files)
    ctx.load(sub_folders)
    ctx.execute_query()

    # 处理当前文件夹内的文件
    print(f"\n当前文件夹: {folder_rel_url}")
    print("文件列表:")
    for file in files:
        print(f"- {file.name}")
        # 直接读取文件内容到内存
        file_content = file.open_binary()
        ctx.execute_query()
        # 示例:转为字符串预览前100字符
        content_str = file_content.content.decode('utf-8')[:100]
        print(f"  内容预览: {content_str}...")

    # 递归处理子文件夹
    print("\n子文件夹列表:")
    for sub_folder in sub_folders:
        print(f"- {sub_folder.name}")
        traverse_folder(sub_folder.serverRelativeUrl)

# 从共享文档根目录开始遍历
traverse_folder("/sites/目标站点名称/Shared Documents")

使用azure-identity + requests调用Graph API

from azure.identity import ClientSecretCredential
import requests

# 配置参数
tenant_id = "你的租户ID"
client_id = "你的AAD应用客户端ID"
client_secret = "你的AAD应用客户端密钥"
site_id = "目标SharePoint站点ID"
drive_id = "站点对应的Drive ID"

# 获取AAD访问令牌
credential = ClientSecretCredential(tenant_id, client_id, client_secret)
token = credential.get_token("https://graph.microsoft.com/.default")
headers = {"Authorization": f"Bearer {token.token}"}

# 递归遍历Drive内的文件和文件夹
def traverse_drive_items(item_id="root"):
    # 获取当前目录下的所有项
    url = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drives/{drive_id}/items/{item_id}/children"
    response = requests.get(url, headers=headers)
    response.raise_for_status()
    items = response.json()["value"]

    for item in items:
        if "folder" in item:
            print(f"\n文件夹: {item['name']}")
            traverse_drive_items(item["id"])
        else:
            print(f"\n文件: {item['name']}")
            # 读取文件内容到内存
            content_url = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drives/{drive_id}/items/{item['id']}/content"
            content_response = requests.get(content_url, headers=headers)
            content_response.raise_for_status()
            # 示例:预览前100字符
            print(f"内容预览: {content_response.text[:100]}...")

# 从Drive根目录开始遍历
traverse_drive_items()

注意事项

  • 权限配置:必须为AAD应用添加应用权限(而非委托权限),并由管理员完成权限同意,否则无法通过应用身份访问SharePoint资源。
  • 站点/Drive ID获取:可通过Graph API调用GET https://graph.microsoft.com/v1.0/sites/你的租户.sharepoint.com:/sites/目标站点名称获取站点ID,再从返回结果的drive字段提取Drive ID。
  • 大文件处理:对于超大文件,可使用API的分块读取功能,依然无需下载到本地,直接在内存中处理分块内容。

内容的提问来源于stack exchange,提问作者bye scraper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 16:15:00