You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Microsoft Graph API读取SharePoint中Doc/Docx文件的文本内容?

问题描述

我在SharePoint中有一个包含n个子目录的目录,每个子目录下都存放着doc或docx格式的文件。我希望读取这些文件的文本内容(转换为纯文本以便解析所有字符串)。我了解docx2txt工具,但它要求文件必须存储在本地机器中。有没有更优的实现方法?目前我正在使用Microsoft Graph API扫描/浏览SharePoint目录,恳请提供相关指导。

当前代码实现
import requests
import pathlib

# 复制获取到的access_token,并指定要调用的MS Graph API端点,例如'https://graph.microsoft.com/v1.0/groups'以获取组织中的所有组
#access_token = '{之前获取的ACCESS TOKEN}'

url = "https://graph.microsoft.com/v1.0/......"
headers = {
  'Authorization': token_result['access_token']
}

consentfilecount=0
clientreportcount = 0
graphlinkcount = 0

while True:
    
    try:
      graph_result = requests.get(url=url, headers=headers)
      graph_result.raise_for_status()
    except:
      token_result = client.acquire_token_for_client(scopes=scope)
    
    headers = {
      'Authorization': token_result['access_token']
    }
    

    if ('value' in graph_result.json()):
      for list in graph_result.json()['value']:
        for ele in finalReportNames:
          if ele.lower() in list["name"].lower():
            clientreportcount +=1
            response = requests.get(list["webUrl"],headers=headers)#{"Authorization": f"Bearer " +token_result['access_token']})
            print(response)
            print(list["name"])
            print(list["webUrl"])
            print(pathlib.Path(list["name"]).suffix)
        #print(graph_result.json())
      if('@odata.nextLink' in graph_result.json()):
        url = graph_result.json()['@odata.nextLink']
        graphlinkcount += 1
      else:
        break

print(consentfilecount)
解决方案

核心思路:无需本地存储,直接通过Graph API+Python库提取文本

不需要将文件下载到本地,结合Microsoft Graph API和Python文档处理库就能直接获取纯文本,具体步骤如下:

  1. 获取文件字节流
    替换代码中请求webUrl的逻辑,改用Graph API的文件内容端点。每个文件项的id可从Graph返回的value数组中获取,构造请求URL:

    https://graph.microsoft.com/v1.0/sites/{site-id}/drive/items/{item-id}/content
    

    发送GET请求即可获取文件的字节流,无需保存到本地。

  2. 处理docx文件
    使用python-docx库直接读取字节流中的文本:

    from docx import Document
    import io
    
    # content为请求获取到的字节流
    doc = Document(io.BytesIO(content))
    full_text = '\n'.join([para.text for para in doc.paragraphs])
    
  3. 处理doc文件
    Graph API支持将doc格式转换为docx,调用转换接口:

    POST https://graph.microsoft.com/v1.0/sites/{site-id}/drive/items/{item-id}/content?format=docx
    

    获取转换后的docx字节流后,再按上述docx的方式提取文本。

代码优化建议

  • 修复token重试逻辑:当前异常处理中仅重新获取token,但未重新发起原请求。正确流程是获取新token后,重新执行请求。
  • 区分文件类型:通过pathlib.Path(item["name"]).suffix判断文件是.doc还是.docx,分别执行对应处理逻辑。
  • 权限验证:确保应用或用户拥有Files.Read或Sites.Read.All权限,否则无法访问文件内容。

修改后的核心代码片段参考:

import requests
import pathlib
from docx import Document
import io

# 初始化参数(需补充实际site-id、客户端信息)
site_id = "你的站点ID"
scope = ["https://graph.microsoft.com/.default"]
client = # 初始化MSAL客户端实例

token_result = client.acquire_token_for_client(scopes=scope)
headers = {'Authorization': token_result['access_token']}
# 递归获取所有doc/docx文件
url = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drive/root/children?recursive=true&$filter=endswith(name,'.doc') or endswith(name,'.docx')"

while True:
    try:
        graph_result = requests.get(url=url, headers=headers)
        graph_result.raise_for_status()
    except requests.exceptions.HTTPError:
        # token过期,重新获取后重试请求
        token_result = client.acquire_token_for_client(scopes=scope)
        headers = {'Authorization': token_result['access_token']}
        graph_result = requests.get(url=url, headers=headers)
        graph_result.raise_for_status()

    if 'value' in graph_result.json():
        for item in graph_result.json()['value']:
            file_suffix = pathlib.Path(item["name"]).suffix.lower()
            if file_suffix == '.docx':
                content_url = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drive/items/{item['id']}/content"
                content_res = requests.get(content_url, headers=headers)
                doc = Document(io.BytesIO(content_res.content))
                text = '\n'.join([para.text for para in doc.paragraphs])
                print(f"文件{item['name']}文本内容:\n{text}")
            elif file_suffix == '.doc':
                convert_url = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drive/items/{item['id']}/content?format=docx"
                convert_res = requests.post(convert_url, headers=headers)
                doc = Document(io.BytesIO(convert_res.content))
                text = '\n'.join([para.text for para in doc.paragraphs])
                print(f"文件{item['name']}文本内容:\n{text}")

        if '@odata.nextLink' in graph_result.json():
            url = graph_result.json()['@odata.nextLink']
        else:
            break

内容的提问来源于stack exchange,提问作者WhoamI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 14:20:55