如何用Python抓取公共SharePoint URL数据?代码报错求助
尝试从SharePoint URL抓取数据并保存到本地Excel文件,但保存的文件不包含在线Excel表格的数据,运行代码时出现报错。
原代码
import requests url = 'https://crokepark-my.sharepoint.com/:x:/r/personal/ruairi_harvey_gaa_ie/_layouts/15/Doc.aspx?guestaccesstoken=Gc2myfwceMcTJO0Sm78dGMt4Up6MH9VlzlUxxsMV%2Fgk%3D&docid=04bc452cba06b4bfea0d1ed80a2b5fac6&action=default&cid=f2e59e9a-71fd-4fd4-aaa6-8e3a4e88fe40' response = requests.get(url) with open('file.HTML', 'wb') as f: f.write(response.content) from bs4 import BeautifulSoup # Open the HTML file with open("file.html", "r") as file: html_content = file.read() # Use BeautifulSoup to parse the HTML soup = BeautifulSoup(html_content, 'html.parser') # Find the table in the HTML using its attributes (e.g. class, id) table = soup.find('table', attrs={'class': 'table-class-name'}) soup # Extract the table headers headers = [header.text for header in table.find_all('Aimn')] # Extract the table data rows rows = [] for row in table.find_all('tr'): rows.append([cell.text for cell in row.find_all('td')]) # Print the table headers and data print(headers) print(rows)
报错信息
--------------------------------------------------------------------------- AttributeError Traceback (most recent call last) ~\\AppData\\Local\\Temp/ipykernel_4532/3438091026.py in <module> 1 # Extract the table headers ----> 2 headers = [header.text for header in table.find_all('Aimn')] 3 4 # Extract the table data rows 5 rows = [] AttributeError: 'NoneType' object has no attribute 'find_all'
问题分析与解决方案
问题根源
- URL错误:你用的是SharePoint在线文档预览页面(
Doc.aspx),该页面通过JavaScript动态渲染表格数据,requests.get只能获取静态HTML源码,拿不到实际表格内容。 - 无效选择器:代码里用
table-class-name作为表格类名查找元素,这是占位符,实际页面不存在该类名,导致table变量为None触发报错;另外find_all('Aimn')也是错误的,HTML中没有名为Aimn的标签。
正确解决方法
直接获取Excel文件的原始下载链接,而非预览页面,修改规则:
- 将原URL中的
/:x:/r/替换为/:x:/d/ - 将末尾的
action=default改为action=download
改进后的代码
import requests import pandas as pd # 修改后的下载URL url = 'https://crokepark-my.sharepoint.com/:x:/d/personal/ruairi_harvey_gaa_ie/_layouts/15/Doc.aspx?guestaccesstoken=Gc2myfwceMcTJO0Sm78dGMt4Up6MH9VlzlUxxsMV%2Fgk%3D&docid=04bc452cba06b4bfea0d1ed80a2b5fac6&action=download&cid=f2e59e9a-71fd-4fd4-aaa6-8e3a4e88fe40' # 发送请求下载文件,允许重定向 response = requests.get(url, allow_redirects=True) # 保存为本地Excel文件 with open('sharepoint_excel.xlsx', 'wb') as f: f.write(response.content) # 读取并查看Excel数据(可选操作) df = pd.read_excel('sharepoint_excel.xlsx') print(df.head())
补充说明
- 你的URL包含
guestaccesstoken,属于访客权限,修改后的下载链接应该可直接访问;如果仍无法下载,需检查令牌有效性或处理额外身份验证。 - 无需再用BeautifulSoup解析HTML,直接下载原始Excel文件是最可靠的方式,避开了动态渲染的问题。
内容的提问来源于stack exchange,提问作者Hedge_hog
相关产品推荐
相关产品推荐

