You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何自动获取受密码保护网站的全新Cookie以实现持续爬虫?

登录态Cookie自动获取与跨库复用方案

一、从Selenium导出Cookie并适配requests格式

Selenium获取的Cookie是列表格式,每个元素包含name、value等字段,需要转换成requests要求的键值对字典才能直接使用:

from selenium import webdriver

# 启动Chrome浏览器(可添加无头模式参数)
driver = webdriver.Chrome()
driver.get("https://www.arbeitsagentur.de/bewerberboerse/")

# 手动完成登录后按回车继续(后续可替换为自动化登录逻辑)
input("登录完成后按回车继续...")

# 提取Selenium中的Cookie列表
selenium_cookies = driver.get_cookies()

# 转换为requests可用的Cookie字典
requests_cookies = {cookie['name']: cookie['value'] for cookie in selenium_cookies}

# 关闭浏览器
driver.quit()

二、解决Selenium登录时元素定位失败问题

针对网站重定向多、结构复杂的情况,用以下方法规避定位失败:

  • 显式等待元素加载:不用固定sleep,等待元素可交互后再操作,避免页面未加载完成导致定位失效
  • 处理iframe嵌套:如果登录框在iframe内,需先切换到对应iframe再定位元素
  • 优先用唯一属性定位:避免依赖绝对XPath,改用元素的id、name或唯一class属性定位

自动化登录示例代码:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 等待登录入口可点击并触发跳转
login_btn = WebDriverWait(driver, 10).until(
    EC.element_to_be_clickable((By.LINK_TEXT, "登录"))
)
login_btn.click()

# 若登录框在iframe中,先切换iframe(替换为实际iframe的id或name)
# driver.switch_to.frame("login-frame")

# 等待用户名输入框加载并输入内容
username_input = WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.ID, "username"))
)
username_input.send_keys("你的用户名")

# 处理密码输入与登录按钮
password_input = driver.find_element(By.ID, "password")
password_input.send_keys("你的密码")
driver.find_element(By.ID, "submit-login").click()

# 等待登录跳转完成(验证首页元素出现)
WebDriverWait(driver, 15).until(
    EC.presence_of_element_located((By.CLASS_NAME, "dashboard-header"))
)

三、整合流程:自动获取Cookie并发起API请求

将Cookie获取与requests请求整合,每次爬取前自动刷新有效Cookie:

import requests

# 先执行上述Selenium代码获取requests_cookies

# 携带Cookie发起API请求
api_url = "目标API接口地址"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

response = requests.get(api_url, cookies=requests_cookies, headers=headers)
if response.status_code == 200:
    json_data = response.json()
    # 处理返回的JSON数据
else:
    print(f"请求失败,状态码:{response.status_code}")

额外注意事项

  • 若网站有反爬机制,可给Selenium添加模拟用户行为(如随机停顿、滚动页面),或使用代理IP
  • Cookie过期后,重新运行Selenium登录流程即可获取新Cookie
  • 可将Cookie保存到本地文件(用json.dump),下次先读取本地Cookie,验证失效后再重新获取

内容的提问来源于stack exchange,提问作者Ali Ahsen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 12:05:56