使用Python获取本地SharePoint页面HTML遇403错误求解决方案
先说说你遇到403错误的可能原因:
- 认证方式不匹配:你的代码用了NTLM认证,但如果你的SharePoint是现代版本(比如基于Microsoft 365的),它可能默认用OAuth2而非NTLM,这时候
HttpNtlmAuth就不起作用了;另外也可能是你输入的域名/账号格式有误,或者账号本身没有访问这个Profile页面的权限(先确认用这个账号在浏览器里能正常打开页面)。 - 请求头缺失:浏览器发送请求时会带一堆标识头(比如User-Agent、Accept等),
requests库默认的请求头太“朴素”,很容易被SharePoint的安全机制拦截。 - 缺少验证令牌:部分SharePoint页面会要求请求携带CSRF令牌或者特定Cookie,直接发送GET请求没有这些验证信息就会被拒绝。
接下来给你几个可行的解决办法:
方法1:完善NTLM认证+补全请求头
如果你的SharePoint确实用NTLM认证,先确认账号权限没问题,然后给请求加上浏览器风格的请求头,用Session保持会话:
import requests from requests_ntlm import HttpNtlmAuth from bs4 import BeautifulSoup # 目标URL url = "https://my.mycompany.net/Profile.aspx?acname=i%3A0%23.f%7Cmembership%7Cparametertext%40company.net" # 模拟浏览器的请求头 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3' } # 创建会话,保持认证状态和请求头 session = requests.Session() session.auth = HttpNtlmAuth('domain\\userid', 'mypassword') session.headers.update(headers) response = session.get(url) print(f"响应状态码:{response.status_code}") if response.status_code == 200: # 解析HTML并提取数据 soup = BeautifulSoup(response.text, "html.parser") # 按类名提取 target_elements = soup.find_all(class_="your-target-class") # 按ID提取 target_element = soup.find(id="your-target-id") # 按标签提取 all_divs = soup.find_all("div")
方法2:使用SharePoint官方API(推荐)
解析HTML其实是比较脆弱的做法,官方API能直接获取结构化数据,还不会被拦截。如果是基于Microsoft 365的SharePoint,推荐用Microsoft Graph API:
import requests # 第一步:获取访问令牌(需要先在Azure AD注册应用,拿到client_id、client_secret、tenant_id) token_url = "https://login.microsoftonline.com/your-tenant-id/oauth2/v2.0/token" token_payload = { "grant_type": "client_credentials", "client_id": "your-client-id", "client_secret": "your-client-secret", "scope": "https://graph.microsoft.com/.default" } token_response = requests.post(token_url, data=token_payload) access_token = token_response.json()["access_token"] # 第二步:请求用户资料 graph_url = "https://graph.microsoft.com/v1.0/users/parametertext@company.net" headers = {"Authorization": f"Bearer {access_token}"} user_response = requests.get(graph_url, headers=headers) if user_response.status_code == 200: user_data = user_response.json() print("用户资料:", user_data) # 直接拿结构化数据,比如用户邮箱、姓名等 print(f"用户名:{user_data['displayName']},邮箱:{user_data['mail']}")
注:注册Azure AD应用的步骤需要你有公司的Azure权限,或者找IT部门协助。
方法3:用Selenium模拟浏览器访问
如果前两种方法都走不通,用Selenium模拟真实浏览器操作,能绕过大部分验证机制:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # 初始化Chrome浏览器(需要提前下载对应版本的ChromeDriver) driver = webdriver.Chrome() driver.get("https://my.mycompany.net/Profile.aspx?acname=i%3A0%23.f%7Cmembership%7Cparametertext%40company.net") # 模拟登录(根据实际登录页面的元素ID调整) driver.find_element(By.ID, "username").send_keys("domain\\userid") driver.find_element(By.ID, "password").send_keys("mypassword") driver.find_element(By.ID, "submit-btn").click() # 等待目标元素加载完成,避免页面未加载就解析 WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.CLASS_NAME, "target-class"))) # 获取页面源码并解析 page_source = driver.page_source soup = BeautifulSoup(page_source, "html.parser") # 提取数据 target_elements = soup.find_all(class_="target-class") for elem in target_elements: print(elem.text) # 关闭浏览器 driver.quit()
内容的提问来源于stack exchange,提问作者IC123456
相关产品推荐
相关产品推荐

