You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

自动化下载psacard.com图片遇PerimeterX拦截问题求助

批量下载PSACard认证卡图片时遭遇PerimeterX拦截问题

需求背景

我需要自动化下载psacard.com上自己送检的评级收藏卡图片,这些图片可在卡片的认证页面获取,示例页面:https://www.psacard.com/cert/68786145/psa。编写代码的目的是高效批量下载对应证书号的图片,用于线上店铺运营,目前手动完成该流程十分繁琐。

问题描述

我已经实现了单张卡片(用户输入单个证书号)的图片下载功能,但将程序扩展为遍历证书号列表时,开始遇到PerimeterX验证码和拦截,而同一浏览器窗口仍可手动访问页面。测试用证书号列表为[68786145,68786146,68786147,68786148],首次运行代码时证书号页面可正常加载,但第二次尝试必然失败。运行脚本时我在操作电脑,鼠标/键盘操作自然,不应被判定为机器人,希望得到解决思路。

现有代码

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from fake_useragent import UserAgent
import time
import random
import requests

certstart= input("Please type the first cert number")
certend= input("Please type in the last cert number")

certs =[]
for i in range(int(certstart), int(certend) + 1):
    certs.append(i)

chrome_options = Options()
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option('useAutomationExtension', False)

driver = webdriver.Chrome(
    executable_path=r'C:\\Users\\danie\\.cache\\selenium\\chromedriver\\win32\\109.0.5414.74\\chromedriver.exe', options=chrome_options)
url = "https://www.psacard.com"
driver.get(url)
time.sleep(3)

for cert in certs:
    driver.implicitly_wait(10)
    url = "https://www.psacard.com/cert/" + str(cert) 
    driver.get(url)
    html_doc = driver.page_source
    print(html_doc)
    time.sleep(2+random.randint(1,4))
    orig_text = html_doc

    first_separator = 'id="psaCertImages"'
    output = []

    orig_text = orig_text.split(" ")
    print (first_separator in orig_text)
    
    if first_separator in orig_text:
        output = orig_text[orig_text.index(first_separator)+1 : ]
        
    imgs = []
    for i in output:
        if "cloudfront" in i:
            if "href" in i:
                url = i.strip("href=")
                url = url.strip('"')
                imgs.append(url)

    print(imgs)
    count = 0
    text = ["front", "back"]
    for i in imgs:
        s = requests.session()
        r = s.get(i, allow_redirects=True)
        open(str(cert) + "-" + text[count] + '.jpeg', 'wb').write(r.content)
        count += 1
        time.sleep(2+random.randint(1,4))
    time.sleep(10 + random.randint(1,10))

内容的提问来源于stack exchange,提问作者parrisimo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 03:21:08