You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python抓取公共SharePoint URL数据?代码报错求助

问题:SharePoint在线Excel数据抓取失败,本地文件无数据

尝试从SharePoint URL抓取数据并保存到本地Excel文件,但保存的文件不包含在线Excel表格的数据,运行代码时出现报错。

原代码

import requests

url = 'https://crokepark-my.sharepoint.com/:x:/r/personal/ruairi_harvey_gaa_ie/_layouts/15/Doc.aspx?guestaccesstoken=Gc2myfwceMcTJO0Sm78dGMt4Up6MH9VlzlUxxsMV%2Fgk%3D&docid=04bc452cba06b4bfea0d1ed80a2b5fac6&action=default&cid=f2e59e9a-71fd-4fd4-aaa6-8e3a4e88fe40'
response = requests.get(url)

with open('file.HTML', 'wb') as f:
    f.write(response.content)

from bs4 import BeautifulSoup

# Open the HTML file
with open("file.html", "r") as file:
    html_content = file.read()

# Use BeautifulSoup to parse the HTML
soup = BeautifulSoup(html_content, 'html.parser')

# Find the table in the HTML using its attributes (e.g. class, id)
table = soup.find('table', attrs={'class': 'table-class-name'})
soup 

# Extract the table headers
headers = [header.text for header in table.find_all('Aimn')]

# Extract the table data rows
rows = []
for row in table.find_all('tr'):
    rows.append([cell.text for cell in row.find_all('td')])

# Print the table headers and data
print(headers)
print(rows)

报错信息

---------------------------------------------------------------------------
AttributeError                            Traceback (most recent call last)
~\\AppData\\Local\\Temp/ipykernel_4532/3438091026.py in <module>
      1 # Extract the table headers
----> 2 headers = [header.text for header in table.find_all('Aimn')]
      3 
      4 # Extract the table data rows
      5 rows = []

AttributeError: 'NoneType' object has no attribute 'find_all'

问题分析与解决方案

问题根源

  1. URL错误:你用的是SharePoint在线文档预览页面(Doc.aspx),该页面通过JavaScript动态渲染表格数据,requests.get只能获取静态HTML源码,拿不到实际表格内容。
  2. 无效选择器:代码里用table-class-name作为表格类名查找元素,这是占位符,实际页面不存在该类名,导致table变量为None触发报错;另外find_all('Aimn')也是错误的,HTML中没有名为Aimn的标签。

正确解决方法

直接获取Excel文件的原始下载链接,而非预览页面,修改规则:

  • 将原URL中的/:x:/r/替换为/:x:/d/
  • 将末尾的action=default改为action=download

改进后的代码

import requests
import pandas as pd

# 修改后的下载URL
url = 'https://crokepark-my.sharepoint.com/:x:/d/personal/ruairi_harvey_gaa_ie/_layouts/15/Doc.aspx?guestaccesstoken=Gc2myfwceMcTJO0Sm78dGMt4Up6MH9VlzlUxxsMV%2Fgk%3D&docid=04bc452cba06b4bfea0d1ed80a2b5fac6&action=download&cid=f2e59e9a-71fd-4fd4-aaa6-8e3a4e88fe40'

# 发送请求下载文件,允许重定向
response = requests.get(url, allow_redirects=True)

# 保存为本地Excel文件
with open('sharepoint_excel.xlsx', 'wb') as f:
    f.write(response.content)

# 读取并查看Excel数据(可选操作)
df = pd.read_excel('sharepoint_excel.xlsx')
print(df.head())

补充说明

  • 你的URL包含guestaccesstoken,属于访客权限,修改后的下载链接应该可直接访问;如果仍无法下载,需检查令牌有效性或处理额外身份验证。
  • 无需再用BeautifulSoup解析HTML,直接下载原始Excel文件是最可靠的方式,避开了动态渲染的问题。

内容的提问来源于stack exchange,提问作者Hedge_hog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 19:31:23