You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从SharePoint加载Excel至Pandas DataFrame遇报错,请求技术协助

问题分析

你遇到的两个错误本质是同一个问题:代码请求的URL返回的不是Excel文件的二进制内容,而是HTML页面(比如SharePoint的文件预览页或未正确跳转的页面),导致Excel解析引擎无法识别内容格式。

解决步骤

1. 获取正确的文件直接访问URL

不要使用浏览器打开文件时地址栏的URL(通常带有?web=1后缀),正确操作:

  • 在SharePoint站点找到目标Excel文件,右键点击→复制链接
  • 若复制的链接末尾带有?web=1,直接删除该后缀;部分版本可直接选择复制「直接下载链接」
  • 正确URL示例:https://yourdomain.sharepoint.com/sites/YourSite/Shared%20Documents/TargetFile.xlsx

2. 修正认证逻辑

AuthenticationContext初始化需传入SharePoint站点的根URL,而非文件URL,否则认证上下文会出现异常。比如站点根URL格式为https://yourdomain.sharepoint.com/sites/YourSite。

修正后的完整代码
from office365.runtime.auth.authentication_context import AuthenticationContext
from office365.sharepoint.client_context import ClientContext
from office365.sharepoint.files.file import File 
import io
import pandas as pd

# 配置参数
site_url = "https://yourdomain.sharepoint.com/sites/YourSite"  # 站点根URL
file_relative_path = "/sites/YourSite/Shared Documents/TargetFile.xlsx"  # 文件相对站点的路径
username = "myname@mydomainemail.co.id"
password = "mypassword"
sheetname = "mytabsheetname"

# 认证流程
ctx_auth = AuthenticationContext(site_url)
if ctx_auth.acquire_token_for_user(username, password):
    ctx = ClientContext(site_url, ctx_auth)
    web = ctx.web
    ctx.load(web)
    ctx.execute_query()
    print("认证成功")
else:
    print(ctx_auth.get_last_error())

# 获取文件二进制内容
response = File.open_binary(ctx, file_relative_path)

# 读取为Pandas DataFrame
bytes_file_obj = io.BytesIO(response.content)
df = pd.read_excel(bytes_file_obj, sheet_name=sheetname, engine="openpyxl")

# 验证数据
print(df.head())
额外注意事项
  • 若你的账号开启了MFA(多因素认证),acquire_token_for_user方式会失效,需改用证书认证或其他MFA兼容的认证方案
  • 确保依赖包为最新版本:pip install --upgrade openpyxl office365-rest-python-client

内容的提问来源于stack exchange,提问作者user12838

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 02:00:25