如何从在线Python服务(如JupyterLite/Google Colab)读取SharePoint文件?
Google Colab 实现方法
方法1:公开共享链接直接下载
如果你的SharePoint文件已经设置了任何人可查看的共享链接,直接用requests就能把文件拉到Colab环境:
import requests from io import BytesIO import pandas as pd # 以CSV/Excel为例,其他格式可自行调整处理逻辑 # 替换成你的SharePoint共享链接,需把末尾的?e=xxx替换为&download=1,并截断多余参数 sharepoint_url = "https://your-sharepoint-site.com/sites/xxx/Shared%20Documents/your-file.csv?e=abc123" download_url = sharepoint_url.replace("?e=", "&download=1").split("&csf=")[0] # 发送请求获取文件 response = requests.get(download_url) response.raise_for_status() # 请求失败直接抛出错误 # 以CSV为例转成DataFrame df = pd.read_csv(BytesIO(response.content)) # 验证导入结果 print(df.head())
方法2:Microsoft Graph API(需权限验证)
如果文件需要登录权限才能访问,用msal库获取令牌后调用Graph API读取:
先安装依赖:
!pip install msal pandas
再编写代码实现:
import msal import requests import pandas as pd from io import BytesIO # 替换为你的租户ID、客户端ID(需提前在Azure AD注册应用) tenant_id = "your-tenant-id" client_id = "your-client-id" scope = ["https://graph.microsoft.com/.default"] username = "你的Office365邮箱账号" password = "你的Office365密码" # 敏感数据建议改用客户端密钥或证书,避免明文暴露 # 获取访问令牌 app = msal.PublicClientApplication(client_id, authority=f"https://login.microsoftonline.com/{tenant_id}") result = app.acquire_token_by_username_password(username, password, scopes=scope) if "access_token" not in result: raise Exception(f"令牌获取失败: {result.get('error_description')}") access_token = result["access_token"] # 替换为目标文件ID(可通过Graph Explorer查询获取) file_id = "your-file-id" graph_url = f"https://graph.microsoft.com/v1.0/drives/items/{file_id}/content" # 请求文件内容 headers = {"Authorization": f"Bearer {access_token}"} response = requests.get(graph_url, headers=headers) response.raise_for_status() # 以Excel为例转成DataFrame df = pd.read_excel(BytesIO(response.content)) print(df.head())
JupyterLite 实现方法
方法1:公开链接 + Pyodide内置工具
JupyterLite基于浏览器运行,需用Pyodide自带的HTTP工具处理跨域请求:
from pyodide.http import open_url import pandas as pd # 替换为公开可访问的SharePoint文件共享链接,需确保链接允许跨域访问 sharepoint_url = "https://your-sharepoint-site.com/sites/xxx/Shared%20Documents/your-file.csv?e=abc123&download=1" # 读取文件内容 content = open_url(sharepoint_url).read() # 以CSV为例转成DataFrame from io import StringIO df = pd.read_csv(StringIO(content)) print(df.head())
注意:如果SharePoint链接存在CORS限制,要么使用公共CORS代理(敏感数据不建议),要么调整文件共享的权限设置。
方法2:手动上传文件(通用但繁琐)
如果文件无法通过链接直接访问,可采用本地下载再上传的方式:
- 在SharePoint中将文件下载到本地电脑
- 在JupyterLite左侧面板点击Upload Files按钮上传文件
- 用Python读取上传后的文件:
import pandas as pd # 替换为你上传的文件名 df = pd.read_csv("your-file.csv") print(df.head())
内容的提问来源于stack exchange,提问作者Michael
相关产品推荐
相关产品推荐

