You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python获取本地SharePoint页面HTML遇403错误求解决方案

解决SharePoint页面请求403错误及数据提取方案

先说说你遇到403错误的可能原因:

  • 认证方式不匹配:你的代码用了NTLM认证,但如果你的SharePoint是现代版本(比如基于Microsoft 365的),它可能默认用OAuth2而非NTLM,这时候HttpNtlmAuth就不起作用了;另外也可能是你输入的域名/账号格式有误,或者账号本身没有访问这个Profile页面的权限(先确认用这个账号在浏览器里能正常打开页面)。
  • 请求头缺失:浏览器发送请求时会带一堆标识头(比如User-Agent、Accept等),requests库默认的请求头太“朴素”,很容易被SharePoint的安全机制拦截。
  • 缺少验证令牌:部分SharePoint页面会要求请求携带CSRF令牌或者特定Cookie,直接发送GET请求没有这些验证信息就会被拒绝。

接下来给你几个可行的解决办法:

方法1:完善NTLM认证+补全请求头

如果你的SharePoint确实用NTLM认证,先确认账号权限没问题,然后给请求加上浏览器风格的请求头,用Session保持会话:

import requests
from requests_ntlm import HttpNtlmAuth
from bs4 import BeautifulSoup

# 目标URL
url = "https://my.mycompany.net/Profile.aspx?acname=i%3A0%23.f%7Cmembership%7Cparametertext%40company.net"

# 模拟浏览器的请求头
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'zh-CN,zh;q=0.8,en-US;q=0.5,en;q=0.3'
}

# 创建会话,保持认证状态和请求头
session = requests.Session()
session.auth = HttpNtlmAuth('domain\\userid', 'mypassword')
session.headers.update(headers)

response = session.get(url)
print(f"响应状态码:{response.status_code}")

if response.status_code == 200:
    # 解析HTML并提取数据
    soup = BeautifulSoup(response.text, "html.parser")
    # 按类名提取
    target_elements = soup.find_all(class_="your-target-class")
    # 按ID提取
    target_element = soup.find(id="your-target-id")
    # 按标签提取
    all_divs = soup.find_all("div")

方法2:使用SharePoint官方API(推荐)

解析HTML其实是比较脆弱的做法,官方API能直接获取结构化数据,还不会被拦截。如果是基于Microsoft 365的SharePoint,推荐用Microsoft Graph API:

import requests

# 第一步:获取访问令牌(需要先在Azure AD注册应用,拿到client_id、client_secret、tenant_id)
token_url = "https://login.microsoftonline.com/your-tenant-id/oauth2/v2.0/token"
token_payload = {
    "grant_type": "client_credentials",
    "client_id": "your-client-id",
    "client_secret": "your-client-secret",
    "scope": "https://graph.microsoft.com/.default"
}

token_response = requests.post(token_url, data=token_payload)
access_token = token_response.json()["access_token"]

# 第二步:请求用户资料
graph_url = "https://graph.microsoft.com/v1.0/users/parametertext@company.net"
headers = {"Authorization": f"Bearer {access_token}"}
user_response = requests.get(graph_url, headers=headers)

if user_response.status_code == 200:
    user_data = user_response.json()
    print("用户资料:", user_data)
    # 直接拿结构化数据,比如用户邮箱、姓名等
    print(f"用户名:{user_data['displayName']},邮箱:{user_data['mail']}")

注:注册Azure AD应用的步骤需要你有公司的Azure权限,或者找IT部门协助。

方法3:用Selenium模拟浏览器访问

如果前两种方法都走不通,用Selenium模拟真实浏览器操作,能绕过大部分验证机制:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# 初始化Chrome浏览器(需要提前下载对应版本的ChromeDriver)
driver = webdriver.Chrome()
driver.get("https://my.mycompany.net/Profile.aspx?acname=i%3A0%23.f%7Cmembership%7Cparametertext%40company.net")

# 模拟登录(根据实际登录页面的元素ID调整)
driver.find_element(By.ID, "username").send_keys("domain\\userid")
driver.find_element(By.ID, "password").send_keys("mypassword")
driver.find_element(By.ID, "submit-btn").click()

# 等待目标元素加载完成,避免页面未加载就解析
WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.CLASS_NAME, "target-class")))

# 获取页面源码并解析
page_source = driver.page_source
soup = BeautifulSoup(page_source, "html.parser")

# 提取数据
target_elements = soup.find_all(class_="target-class")
for elem in target_elements:
    print(elem.text)

# 关闭浏览器
driver.quit()

内容的提问来源于stack exchange,提问作者IC123456

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:04:30