Python 3.7编写Adidas网站监控脚本遇TimeoutError问题求助
嘿,我看你这脚本遇到了连接超时的麻烦,这在爬取Adidas这类有反爬机制的电商站点时太常见了。先看看你的代码和报错,咱们一步步来搞定它:
你的原始代码
# Import requests (to download the page) import requests # Import BeautifulSoup (to parse what we download) from bs4 import BeautifulSoup # Import Time (to add a delay between the times the scape runs) import time # Import smtplib (to allow us to email) import smtplib # This is a pretty simple script. The script downloads the homepage of VentureBeat, and if it finds some text, emails me. # If it does not find some text, it waits 60 seconds and downloads the homepage again. # while this is true (it is true by default), while True: # set the url as the page i want to monitor, url = "https://www.adidas.co.uk/yeezy" # set the headers like we are a browser, headers = {'user-agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36'} # download the homepage response = requests.get(url, headers=headers) # parse the downloaded homepage and grab all text, then, soup = BeautifulSoup(response.text, "lxml") # if the number of times the word "SELECT SIZE" occurs on the page is less than 1, if str(soup).find("YEEZY BOOST 350 V2 ADULTS") == -1: # wait 60 seconds, time.sleep(60) # continue with the script, continue # but if the word occurs any other number of times, else: gmail_user = 'example@gmail.com' gmail_password = 'Password' sent_from = gmail_user to = ['example1@gmail.com'] subject = 'OMG Super Important Message' body = 'Hey, what up?\n\n- You' email_text = """\ From: %s To: %s Subject: %s %s """ % (sent_from, ", ".join(to), subject, body) try: server = smtplib.SMTP_SSL('smtp.gmail.com', 465) server.ehlo() server.login(gmail_user, gmail_password) server.sendmail(sent_from, to, email_text) server.close() print ('Email sent!') except: print ('Something went wrong...') break
报错信息
Traceback (most recent call last):
...
requests.exceptions.ConnectionError: ('Connection aborted.', TimeoutError(10060, 'A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond', None, 10060, None))
TimeoutError: [WinError 10060] 连接尝试失败,因为连接方在一段时间后未正确响应,或已建立的连接失败,因为连接的主机未能响应
问题分析与解决方案
1. 给请求添加超时时间
你的requests.get没设置超时参数,一旦Adidas服务器迟迟不响应,脚本会一直挂到系统强制中断。给请求加个10秒超时,能及时捕获异常:
response = requests.get(url, headers=headers, timeout=10)
2. 增加请求异常处理,避免脚本崩溃
在请求环节加try-except块,捕获超时、连接失败、HTTP错误等异常,这样脚本不会直接挂掉,而是能继续重试:
try: response = requests.get(url, headers=headers, timeout=10) response.raise_for_status() # 检查请求是否返回4xx/5xx错误 except (requests.exceptions.Timeout, requests.exceptions.ConnectionError, requests.exceptions.HTTPError) as e: print(f"请求出错: {e}") time.sleep(60) continue
3. 优化Headers,模拟真实浏览器
你用的User-Agent版本太老了,Adidas很容易识别出这是爬虫。换成最新的浏览器UA,再补充几个常用Headers,让请求更像真人操作:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.adidas.co.uk/' }
4. 用随机等待时间替代固定延迟
固定60秒等待太规律,容易被反爬系统检测到。用random模块加一点波动,比如在50-70秒之间随机等待:
import random # 替换原来的time.sleep(60) time.sleep(random.randint(50, 70))
5. 可选:使用代理IP轮换
如果还是频繁超时,大概率是你的IP被Adidas限制了。可以找一些靠谱的代理IP,在请求时加入代理参数:
proxies = { 'http': 'http://your-proxy-ip:port', 'https': 'https://your-proxy-ip:port' } # 在requests.get中添加proxies参数 response = requests.get(url, headers=headers, timeout=10, proxies=proxies)
修改后的完整代码
import requests from bs4 import BeautifulSoup import time import smtplib import random while True: url = "https://www.adidas.co.uk/yeezy" headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.adidas.co.uk/' } try: response = requests.get(url, headers=headers, timeout=10) response.raise_for_status() except (requests.exceptions.Timeout, requests.exceptions.ConnectionError, requests.exceptions.HTTPError) as e: print(f"请求失败,稍后重试: {e}") time.sleep(random.randint(50, 70)) continue soup = BeautifulSoup(response.text, "lxml") if "YEEZY BOOST 350 V2 ADULTS" not in str(soup): print("未找到目标文本,等待重试...") time.sleep(random.randint(50, 70)) continue else: gmail_user = 'example@gmail.com' gmail_password = 'Password' sent_from = gmail_user to = ['example1@gmail.com'] subject = 'Yeezy监控触发!' body = '检测到YEEZY BOOST 350 V2 ADULTS页面更新,快去看看!\n\n- 监控脚本' email_text = """\ From: %s To: %s Subject: %s %s """ % (sent_from, ", ".join(to), subject, body) try: server = smtplib.SMTP_SSL('smtp.gmail.com', 465) server.ehlo() server.login(gmail_user, gmail_password) server.sendmail(sent_from, to, email_text) server.close() print ('通知邮件已发送!') except Exception as e: print (f'邮件发送失败: {e}') break
修改后脚本的稳定性会提升不少,也能更好地绕过Adidas的反爬机制。记得把邮箱账号密码换成你自己的,如果用Gmail的话,可能需要开启App Password或者降低应用权限限制。
内容的提问来源于stack exchange,提问作者Gmt_344--

