You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

沙特Etimad招标网站网页爬取失败问题求助

沙特Etimad招标网站网页爬取失败问题求助

看起来你遇到了反爬机制的阻碍啦,别着急,咱们一步步拆解问题、找解决方案:

问题根源分析

你用requests直接请求目标网站失败,但ScrapeStorm能成功,核心原因在于:

  • 你的requests请求太“简洁”,没有模拟真实浏览器的请求特征,被网站的反爬系统识别为爬虫拦截了;
  • ScrapeStorm这类专业爬虫工具默认会模拟浏览器的请求头、会话甚至动态渲染行为,所以能绕过基础的反爬检测。

具体解决步骤

1. 先给请求添加浏览器请求头

网站通常会通过User-Agent、Accept等请求头字段判断请求是否来自真实浏览器。你可以按以下方式修改代码:

import requests
from bs4 import BeautifulSoup

# 从浏览器开发者工具复制真实的请求头(F12 -> Network -> 选中目标请求 -> Headers)
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Connection': 'keep-alive',
    'Upgrade-Insecure-Requests': '1'
}

# 带请求头发送请求
result = requests.get("https://tenders.etimad.sa/Qualification/QualificationsForVisitor", headers=headers)
src = result.content
soup = BeautifulSoup(src,"lxml")
print(soup)

2. 用Session保持会话(如果需要Cookie验证)

有些网站会在首次访问时设置会话Cookie,后续请求需要携带Cookie才能正常访问。可以用requests.Session来自动维护会话:

import requests
from bs4 import BeautifulSoup

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    # 其他必要请求头...
}

session = requests.Session()
# 先访问网站首页获取会话Cookie
session.get("https://tenders.etimad.sa/", headers=headers)
# 再请求目标页面
result = session.get("https://tenders.etimad.sa/Qualification/QualificationsForVisitor", headers=headers)
src = result.content
soup = BeautifulSoup(src,"lxml")
print(soup)

3. 如果是动态渲染内容,改用无头浏览器工具

如果网站内容是通过JavaScript动态加载的,requests只能获取静态HTML,这时候需要用Selenium或Playwright这类工具模拟浏览器渲染:
以Selenium为例:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup

options = Options()
options.add_argument('--headless=new')  # 无头模式,不弹出浏览器窗口
options.add_argument('user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36')

# 初始化浏览器驱动(需要提前下载对应版本的ChromeDriver)
driver = webdriver.Chrome(options=options)
driver.get("https://tenders.etimad.sa/Qualification/QualificationsForVisitor")
# 获取渲染后的页面源码
src = driver.page_source
soup = BeautifulSoup(src,"lxml")
print(soup)
# 关闭浏览器
driver.quit()

总结

你的原始代码没有模拟真实浏览器的请求特征,所以被反爬机制拦截了。建议先从添加请求头开始尝试,这是解决大部分基础反爬问题的第一步;如果还是不行,再考虑会话保持或动态渲染的方案。

备注:内容来源于stack exchange,提问作者mohamed sultan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.21 13:38:10