You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取Angel.co遇HTTP 403错误,求解决方案及翻页爬取方法

解决Angel.co爬取的403错误及分页问题

一、修复requests请求的403错误

403是网站反爬拦截的结果,无法直接“转为200”,需通过完善请求参数绕过检测。原有代码存在两个核心问题:首次请求未携带User-Agent导致获取无效Cookies,且缺少其他必要请求头。

改进后的代码示例:

import requests
import random

user_agents_list = [
    'Mozilla/5.0 (iPad; CPU OS 12_2 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Mobile/15E148',
    'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.83 Safari/537.36',
    'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.51 Safari/537.36',
    'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/105.0.0.0 Safari/537.36',
]

# 构造完整请求头,模拟真实浏览器行为
headers = {
    'User-Agent': random.choice(user_agents_list),
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://angel.co/'
}

session = requests.Session()
# 先访问主页获取有效Cookies
session.get('https://angel.co/', headers=headers)
# 再请求目标页面
response = session.get('https://angel.co/startups', headers=headers)

print(f"状态码: {response.status_code}")
print("页面内容片段:", response.text[:500])

二、cloudscraper实现分页爬取

Angel.co的分页URL格式为https://angel.co/startups?page=N(N为页码,范围1-2500),通过循环遍历页码即可批量爬取:

import cloudscraper
import random
import time

scraper = cloudscraper.create_scraper()

# 遍历所有页码
for page_num in range(1, 2501):
    target_url = f'https://angel.co/startups?page={page_num}'
    response = scraper.get(target_url)
    
    if response.status_code == 200:
        # 保存页面内容到文件,按需替换处理逻辑
        with open(f'angel_startups_page_{page_num}.html', 'w', encoding='utf-8') as file:
            file.write(response.text)
        print(f"第{page_num}页爬取完成")
    else:
        print(f"第{page_num}页爬取失败,状态码: {response.status_code}")
    
    # 添加随机延迟,降低反爬触发概率
    time.sleep(random.uniform(1, 3))

注意事项

  • 控制请求频率,避免短时间内发送大量请求导致IP封禁
  • 若再次出现403,可尝试更换User-Agent池或使用代理IP
  • 爬取前请确认网站robots.txt规则,遵守爬虫协议

内容的提问来源于stack exchange,提问作者vish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 04:10:40