Heroku部署Python请求代理应用遇异常:误判网站收录状态
谷歌收录检测工具问题排查与解决方案
工具概述
我基于Python requests模块开发了一款检测网站是否被谷歌收录的应用,核心逻辑是爬取谷歌搜索页面 www.google.com/search?site:{url}&num=3,通过匹配页面特定短语判断网站未收录状态。
检查逻辑代码
# checking logic response = self.proxy_request(INDEXING_SEARCH_STRING.format(current_url)) if response.status_code != 200: return current_url, False, "failed" soup = bs4.BeautifulSoup(response.text, "html.parser") not_indexed_regex = re.compile("did not match any documents") if soup(text=not_indexed_regex): return current_url, False, "checked" else: print(response.text) return current_url, True, "checked"
代理请求代码
# proxy requests def proxy_request(self, url, **kwargs): fail_count = 0 max_failures = 3 # Adjust this threshold as needed print("Evaluating: ", self.url_manager.current_url_index, "URL: ", url) while fail_count < max_failures: current_proxy = self.proxy_manager.get_proxy_for_request() if current_proxy is None: ProgressManager.update_progress("All given proxy failed") return requests.get(url, **kwargs) try: response = requests.get(url, proxies=current_proxy, timeout=20) if response.status_code == 200: print("Success!") self.proxy_manager.update_proxy() return response else: print("Failed!",response.status_code) ProgressManager.update_progress("Proxy failing with status code: " + str(response.status_code)) time.sleep(0.5) self.proxy_manager.update_proxy() except Exception as e: print("Failed!", e) fail_count += 1 self.proxy_manager.update_proxy() ProgressManager.update_progress(f"Request failed! {e.__class__.__name__}. ") break time.sleep(5) return requests.get(url,timeout=20)
问题1:Heroku代理环境下误判未收录网站
问题现象
本地环境使用/不使用代理均正常,但部署到Heroku后,使用代理时部分未被收录的网站被误判为已收录(返回True,"checked"),不使用代理时结果正确。
原因分析
谷歌会根据访问IP的地区返回不同语言或表述的页面内容,代理IP对应的地区可能未返回"did not match any documents"这个英文短语,导致正则匹配失效。
解决方案
- 扩展匹配关键词:加入多语言/多表述的未收录提示,比如英文的"no results found"、"not found"等,覆盖不同地区的页面内容。
- 基于页面结构判断:谷歌未收录时通常没有搜索结果容器(class为
g的div),结合结构判断提高准确性:# 修改检查逻辑 soup = bs4.BeautifulSoup(response.text, "html.parser") # 先检查搜索结果容器 search_results = soup.find_all("div", class_="g") not_indexed_regex = re.compile(r"did not match any documents|no results found|not found") # 双重验证:无结果容器 或 匹配到未收录文本 if not search_results or soup(text=not_indexed_regex): return current_url, False, "checked" else: return current_url, True, "checked" - 查看Heroku日志:打印代理返回的页面内容到日志,确认实际未收录提示文本,针对性调整正则。
问题2:Heroku长运行进程H-12超时错误(无需额外服务器)
问题原因
Heroku Web Dyno默认请求超时时间为30秒,长运行的检测任务会触发H-12错误。
解决方案
- 拆分任务到Worker Dyno:
- 将Web Dyno仅用于接收任务请求、返回任务状态,实际检测逻辑放到Worker Dyno中运行(Heroku免费额度包含1个Worker)。
- 使用轻量任务队列(如
apscheduler)在Worker中批量处理URL,避免单次任务运行时间过长。
- 优化任务粒度:
- 将待检测URL分成小批量,每次处理10-20个,处理完一批再启动下一批,避免单次进程持续超时。
- 异步请求优化:
- 使用
aiohttp或requests-futures实现异步请求,提高检测效率,减少单批任务的运行时间。
- 使用
问题3:代理401认证失败/IP黑名单错误
问题现象
出现报错:HTTPSConnectionPool(host='www.google.com', port=443): Max retries exceeded with url: /search?q=site:{{URL}}/&num=1 (Caused by ProxyError('Cannot connect to proxy.', OSError('Tunnel connection failed: 401 Auth Failed ip_blacklisted: 3.85.57.0/24')))
解决方案
- 校验代理认证信息:确保代理字典包含正确的用户名和密码,格式应为:
current_proxy = { "http": "http://username:password@proxy_ip:port", "https": "https://username:password@proxy_ip:port" } - 过滤失效代理:
- 在
proxy_request的异常处理中,针对401或IP黑名单错误,标记该代理为失效,不再复用,并继续尝试下一个代理:except Exception as e: print("Failed!", e) fail_count += 1 # 针对401或IP黑名单做特殊处理 if "401 Auth Failed" in str(e) or "ip_blacklisted" in str(e): self.proxy_manager.mark_proxy_as_bad(current_proxy) self.proxy_manager.update_proxy() ProgressManager.update_progress(f"Request failed! {e.__class__.__name__}. ") # 移除break,继续尝试下一个代理 # break
- 在
- 更换高质量代理池:使用定期更新、不易被谷歌拉黑的代理服务,避免使用免费或低质量代理。
内容的提问来源于stack exchange,提问作者Sujal Choudhari
相关产品推荐
相关产品推荐

