如何通过Microsoft Graph API读取SharePoint中Doc/Docx文件的文本内容?
问题描述
我在SharePoint中有一个包含n个子目录的目录,每个子目录下都存放着doc或docx格式的文件。我希望读取这些文件的文本内容(转换为纯文本以便解析所有字符串)。我了解docx2txt工具,但它要求文件必须存储在本地机器中。有没有更优的实现方法?目前我正在使用Microsoft Graph API扫描/浏览SharePoint目录,恳请提供相关指导。
当前代码实现
import requests import pathlib # 复制获取到的access_token,并指定要调用的MS Graph API端点,例如'https://graph.microsoft.com/v1.0/groups'以获取组织中的所有组 #access_token = '{之前获取的ACCESS TOKEN}' url = "https://graph.microsoft.com/v1.0/......" headers = { 'Authorization': token_result['access_token'] } consentfilecount=0 clientreportcount = 0 graphlinkcount = 0 while True: try: graph_result = requests.get(url=url, headers=headers) graph_result.raise_for_status() except: token_result = client.acquire_token_for_client(scopes=scope) headers = { 'Authorization': token_result['access_token'] } if ('value' in graph_result.json()): for list in graph_result.json()['value']: for ele in finalReportNames: if ele.lower() in list["name"].lower(): clientreportcount +=1 response = requests.get(list["webUrl"],headers=headers)#{"Authorization": f"Bearer " +token_result['access_token']}) print(response) print(list["name"]) print(list["webUrl"]) print(pathlib.Path(list["name"]).suffix) #print(graph_result.json()) if('@odata.nextLink' in graph_result.json()): url = graph_result.json()['@odata.nextLink'] graphlinkcount += 1 else: break print(consentfilecount)
解决方案
核心思路:无需本地存储,直接通过Graph API+Python库提取文本
不需要将文件下载到本地,结合Microsoft Graph API和Python文档处理库就能直接获取纯文本,具体步骤如下:
获取文件字节流
替换代码中请求webUrl的逻辑,改用Graph API的文件内容端点。每个文件项的id可从Graph返回的value数组中获取,构造请求URL:https://graph.microsoft.com/v1.0/sites/{site-id}/drive/items/{item-id}/content发送GET请求即可获取文件的字节流,无需保存到本地。
处理docx文件
使用python-docx库直接读取字节流中的文本:from docx import Document import io # content为请求获取到的字节流 doc = Document(io.BytesIO(content)) full_text = '\n'.join([para.text for para in doc.paragraphs])处理doc文件
Graph API支持将doc格式转换为docx,调用转换接口:POST https://graph.microsoft.com/v1.0/sites/{site-id}/drive/items/{item-id}/content?format=docx获取转换后的docx字节流后,再按上述docx的方式提取文本。
代码优化建议
- 修复token重试逻辑:当前异常处理中仅重新获取token,但未重新发起原请求。正确流程是获取新token后,重新执行请求。
- 区分文件类型:通过
pathlib.Path(item["name"]).suffix判断文件是.doc还是.docx,分别执行对应处理逻辑。 - 权限验证:确保应用或用户拥有
Files.Read或Sites.Read.All权限,否则无法访问文件内容。
修改后的核心代码片段参考:
import requests import pathlib from docx import Document import io # 初始化参数(需补充实际site-id、客户端信息) site_id = "你的站点ID" scope = ["https://graph.microsoft.com/.default"] client = # 初始化MSAL客户端实例 token_result = client.acquire_token_for_client(scopes=scope) headers = {'Authorization': token_result['access_token']} # 递归获取所有doc/docx文件 url = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drive/root/children?recursive=true&$filter=endswith(name,'.doc') or endswith(name,'.docx')" while True: try: graph_result = requests.get(url=url, headers=headers) graph_result.raise_for_status() except requests.exceptions.HTTPError: # token过期,重新获取后重试请求 token_result = client.acquire_token_for_client(scopes=scope) headers = {'Authorization': token_result['access_token']} graph_result = requests.get(url=url, headers=headers) graph_result.raise_for_status() if 'value' in graph_result.json(): for item in graph_result.json()['value']: file_suffix = pathlib.Path(item["name"]).suffix.lower() if file_suffix == '.docx': content_url = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drive/items/{item['id']}/content" content_res = requests.get(content_url, headers=headers) doc = Document(io.BytesIO(content_res.content)) text = '\n'.join([para.text for para in doc.paragraphs]) print(f"文件{item['name']}文本内容:\n{text}") elif file_suffix == '.doc': convert_url = f"https://graph.microsoft.com/v1.0/sites/{site_id}/drive/items/{item['id']}/content?format=docx" convert_res = requests.post(convert_url, headers=headers) doc = Document(io.BytesIO(convert_res.content)) text = '\n'.join([para.text for para in doc.paragraphs]) print(f"文件{item['name']}文本内容:\n{text}") if '@odata.nextLink' in graph_result.json(): url = graph_result.json()['@odata.nextLink'] else: break
内容的提问来源于stack exchange,提问作者WhoamI
相关产品推荐
相关产品推荐

