You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从在线Python服务(如JupyterLite/Google Colab)读取SharePoint文件?

从SharePoint读取文件并导入在线Python服务(Google Colab/JupyterLite)

Google Colab 实现方法

方法1:公开共享链接直接下载

如果你的SharePoint文件已经设置了任何人可查看的共享链接,直接用requests就能把文件拉到Colab环境:

import requests
from io import BytesIO
import pandas as pd  # 以CSV/Excel为例,其他格式可自行调整处理逻辑

# 替换成你的SharePoint共享链接,需把末尾的?e=xxx替换为&download=1,并截断多余参数
sharepoint_url = "https://your-sharepoint-site.com/sites/xxx/Shared%20Documents/your-file.csv?e=abc123"
download_url = sharepoint_url.replace("?e=", "&download=1").split("&csf=")[0]

# 发送请求获取文件
response = requests.get(download_url)
response.raise_for_status()  # 请求失败直接抛出错误

# 以CSV为例转成DataFrame
df = pd.read_csv(BytesIO(response.content))
# 验证导入结果
print(df.head())

方法2:Microsoft Graph API(需权限验证)

如果文件需要登录权限才能访问,用msal库获取令牌后调用Graph API读取:

先安装依赖:

!pip install msal pandas

再编写代码实现:

import msal
import requests
import pandas as pd
from io import BytesIO

# 替换为你的租户ID、客户端ID(需提前在Azure AD注册应用)
tenant_id = "your-tenant-id"
client_id = "your-client-id"
scope = ["https://graph.microsoft.com/.default"]
username = "你的Office365邮箱账号"
password = "你的Office365密码"  # 敏感数据建议改用客户端密钥或证书,避免明文暴露

# 获取访问令牌
app = msal.PublicClientApplication(client_id, authority=f"https://login.microsoftonline.com/{tenant_id}")
result = app.acquire_token_by_username_password(username, password, scopes=scope)

if "access_token" not in result:
    raise Exception(f"令牌获取失败: {result.get('error_description')}")

access_token = result["access_token"]

# 替换为目标文件ID(可通过Graph Explorer查询获取)
file_id = "your-file-id"
graph_url = f"https://graph.microsoft.com/v1.0/drives/items/{file_id}/content"

# 请求文件内容
headers = {"Authorization": f"Bearer {access_token}"}
response = requests.get(graph_url, headers=headers)
response.raise_for_status()

# 以Excel为例转成DataFrame
df = pd.read_excel(BytesIO(response.content))
print(df.head())

JupyterLite 实现方法

方法1:公开链接 + Pyodide内置工具

JupyterLite基于浏览器运行,需用Pyodide自带的HTTP工具处理跨域请求:

from pyodide.http import open_url
import pandas as pd

# 替换为公开可访问的SharePoint文件共享链接,需确保链接允许跨域访问
sharepoint_url = "https://your-sharepoint-site.com/sites/xxx/Shared%20Documents/your-file.csv?e=abc123&download=1"

# 读取文件内容
content = open_url(sharepoint_url).read()

# 以CSV为例转成DataFrame
from io import StringIO
df = pd.read_csv(StringIO(content))
print(df.head())

注意:如果SharePoint链接存在CORS限制,要么使用公共CORS代理(敏感数据不建议),要么调整文件共享的权限设置。

方法2:手动上传文件(通用但繁琐)

如果文件无法通过链接直接访问,可采用本地下载再上传的方式:

  1. 在SharePoint中将文件下载到本地电脑
  2. 在JupyterLite左侧面板点击Upload Files按钮上传文件
  3. 用Python读取上传后的文件:
import pandas as pd

# 替换为你上传的文件名
df = pd.read_csv("your-file.csv")
print(df.head())

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 10:02:56