Scrapy启动附件下载时触发TCP连接超时错误的排查求助
Scrapy附件下载时出现TCP连接超时错误排查
问题背景
使用Scrapy从目标网站提取数据,流程为主页面进入子页面,子页面提取数据并下载附件到指定文件夹。数据提取功能正常,但启动附件下载就触发TCP连接超时错误,已尝试设置DOWNLOAD_DELAY但无效。
错误日志
[scrapy.downloadermiddlewares.retry] DEBUG: Retrying <GET http://59.207.152.10:8001/robots.txt> (failed 1 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond.. 2023-06-06 04:23:41 [scrapy.downloadermiddlewares.retry] DEBUG: Retrying <GET http://59.207.152.10:8001/robots.txt> (failed 2 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond.. 2023-06-06 04:23:56 [scrapy.extensions.logstats] INFO: Crawled 24 pages (at 24 pages/min), scraped 0 items (at 0 items/min) 2023-06-06 04:24:02 [scrapy.downloadermiddlewares.retry] ERROR: Gave up retrying <GET http://59.207.152.10:8001/robots.txt> (failed 3 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond.. 2023-06-06 04:24:02 [scrapy.downloadermiddlewares.robotstxt] ERROR: Error downloading <GET http://59.207.152.10:8001/robots.txt>: TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond.. Traceback (most recent call last): File "C:\Users\Mudassir\AppData\Local\Programs\Python\Python311\Lib\site-packages\scrapy\core\downloader\middleware.py", line 54, in process_request return (yield download_func(request=request, spider=spider)) twisted.internet.error.TCPTimedOutError: TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond.. 2023-06-06 04:24:24 [scrapy.downloadermiddlewares.retry] DEBUG: Retrying <GET http://59.207.152.10:8001/pdshbj/upload/files/2023/5/ab174cf3fe6d1e68dc82bf6ba46a365f.doc> (failed 1 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond.. 2023-06-06 04:24:45 [scrapy.downloadermiddlewares.retry] DEBUG: Retrying <GET http://59.207.152.10:8001/pdshbj/upload/files/2023/5/ab174cf3fe6d1e68dc82bf6ba46a365f.doc> (failed 2 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond.. 2023-06-06 04:24:56 [scrapy.extensions.logstats] INFO: Crawled 24 pages (at 0 pages/min), scraped 0 items (at 0 items/min) 2023-06-06 04:25:06 [scrapy.downloadermiddlewares.retry] ERROR: Gave up retrying <GET http://59.207.152.10:8001/pdshbj/upload/files/2023/5/ab174cf3fe6d1e68dc82bf6ba46a365f.doc> (failed 3 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond.. 2023-06-06 04:25:06 [scrapy.core.scraper] ERROR: Error downloading <GET http://59.207.152.10:8001/pdshbj/upload/files/2023/5/ab174cf3fe6d1e68dc82bf6ba46a365f.doc> Traceback (most recent call last): File "C:\Users\Mudassir\AppData\Local\Programs\Python\Python311\Lib\site-packages\scrapy\core\downloader\middleware.py", line 54, in process_request return (yield download_func(request=request, spider=spider)) twisted.internet.error.TCPTimedOutError: TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond..
爬虫代码(已移除文本提取部分)
class SthjjSpiders(scrapy.Spider): name="sthjj" start_urls=[ "xyz.html" ] def parse(self, response): links=response.css('div.xxgk td a::attr(href)').getall() for link in links: yield response.follow(link, self.page_parse) next_page= response.css('span.item.operation > a::attr(href)')[2].get() self.project_count+=1 if next_page and self.project_count<1: yield response.follow(next_page, callback=self.parse) def page_parse(self, response): attachments=response.css('a[href*=".doc"]::attr(href)').getall() attachments_folder=os.path.join(page_dir,'Attachments') os.makedirs(attachments_folder, exist_ok=True) for link in attachments: yield response.follow(link, callback=self.download_attachments, meta={'attachments_folder': attachments_folder}) def download_attachments(self, response): filename=response.url.split('/')[-1] attachment_folder=response.meta['attachments_folder'] attachments_file=os.path.join(attachment_folder,filename) with open(attachments_file, 'wb') as f: f.write(response.body)
可能的原因
- 目标服务器限制:目标服务器针对大文件(如DOC附件)的下载请求设置了更严格的访问限制,比如IP频率拦截、连接数限制,或者服务器本身带宽不足,处理附件请求时响应缓慢。
- 网络层面拦截:本地网络的防火墙、代理设置,或者目标服务器所在的网络屏蔽了你的IP,导致无法建立TCP连接。可直接在浏览器中访问附件URL,排查是否能正常下载。
- Scrapy下载器配置不足:仅设置
DOWNLOAD_DELAY不够,需调整DOWNLOAD_TIMEOUT延长超时时间,或降低CONCURRENT_REQUESTS减少并发请求数,避免触发服务器防御机制。 - 服务器状态异常:目标服务器的文件存储节点出现故障,导致无法响应附件下载请求。
- URL拼接问题:
response.follow可能拼接出错误的附件URL,可打印实际请求的URL确认是否正确。
内容的提问来源于stack exchange,提问作者Mudassir Ahmed
相关产品推荐
相关产品推荐

