You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy启动附件下载时触发TCP连接超时错误的排查求助

Scrapy附件下载时出现TCP连接超时错误排查

问题背景

使用Scrapy从目标网站提取数据,流程为主页面进入子页面,子页面提取数据并下载附件到指定文件夹。数据提取功能正常,但启动附件下载就触发TCP连接超时错误,已尝试设置DOWNLOAD_DELAY但无效。

错误日志

[scrapy.downloadermiddlewares.retry] DEBUG: Retrying <GET http://59.207.152.10:8001/robots.txt> (failed 1 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond..
2023-06-06 04:23:41 [scrapy.downloadermiddlewares.retry] DEBUG: Retrying <GET http://59.207.152.10:8001/robots.txt> (failed 2 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond..
2023-06-06 04:23:56 [scrapy.extensions.logstats] INFO: Crawled 24 pages (at 24 pages/min), scraped 0 items (at 0 items/min)
2023-06-06 04:24:02 [scrapy.downloadermiddlewares.retry] ERROR: Gave up retrying <GET http://59.207.152.10:8001/robots.txt> (failed 3 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond..
2023-06-06 04:24:02 [scrapy.downloadermiddlewares.robotstxt] ERROR: Error downloading <GET http://59.207.152.10:8001/robots.txt>: TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond..
Traceback (most recent call last):
  File "C:\Users\Mudassir\AppData\Local\Programs\Python\Python311\Lib\site-packages\scrapy\core\downloader\middleware.py", line 54, in process_request
    return (yield download_func(request=request, spider=spider))
twisted.internet.error.TCPTimedOutError: TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond..
2023-06-06 04:24:24 [scrapy.downloadermiddlewares.retry] DEBUG: Retrying <GET http://59.207.152.10:8001/pdshbj/upload/files/2023/5/ab174cf3fe6d1e68dc82bf6ba46a365f.doc> (failed 1 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond..
2023-06-06 04:24:45 [scrapy.downloadermiddlewares.retry] DEBUG: Retrying <GET http://59.207.152.10:8001/pdshbj/upload/files/2023/5/ab174cf3fe6d1e68dc82bf6ba46a365f.doc> (failed 2 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond..
2023-06-06 04:24:56 [scrapy.extensions.logstats] INFO: Crawled 24 pages (at 0 pages/min), scraped 0 items (at 0 items/min)
2023-06-06 04:25:06 [scrapy.downloadermiddlewares.retry] ERROR: Gave up retrying <GET http://59.207.152.10:8001/pdshbj/upload/files/2023/5/ab174cf3fe6d1e68dc82bf6ba46a365f.doc> (failed 3 times): TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond..
2023-06-06 04:25:06 [scrapy.core.scraper] ERROR: Error downloading <GET http://59.207.152.10:8001/pdshbj/upload/files/2023/5/ab174cf3fe6d1e68dc82bf6ba46a365f.doc>
Traceback (most recent call last):
  File "C:\Users\Mudassir\AppData\Local\Programs\Python\Python311\Lib\site-packages\scrapy\core\downloader\middleware.py", line 54, in process_request
    return (yield download_func(request=request, spider=spider))
twisted.internet.error.TCPTimedOutError: TCP connection timed out: 10060: A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond..

爬虫代码(已移除文本提取部分)

class SthjjSpiders(scrapy.Spider):
    name="sthjj"
    start_urls=[
        "xyz.html"
    ]

    def parse(self, response):
        links=response.css('div.xxgk td a::attr(href)').getall()
        for link in links:
            yield response.follow(link, self.page_parse)
                    
        
        next_page= response.css('span.item.operation > a::attr(href)')[2].get()
        self.project_count+=1
        if next_page and self.project_count<1:
            yield response.follow(next_page, callback=self.parse)


    def page_parse(self, response):
       
        attachments=response.css('a[href*=".doc"]::attr(href)').getall()
        attachments_folder=os.path.join(page_dir,'Attachments')
        os.makedirs(attachments_folder, exist_ok=True)

        for link in attachments:
            yield response.follow(link, callback=self.download_attachments, meta={'attachments_folder': attachments_folder})
            
               
    def download_attachments(self, response):
        
        filename=response.url.split('/')[-1]
        attachment_folder=response.meta['attachments_folder']
        attachments_file=os.path.join(attachment_folder,filename)
        
        
        with open(attachments_file, 'wb') as f:
                f.write(response.body)

可能的原因

  • 目标服务器限制:目标服务器针对大文件(如DOC附件)的下载请求设置了更严格的访问限制,比如IP频率拦截、连接数限制,或者服务器本身带宽不足,处理附件请求时响应缓慢。
  • 网络层面拦截:本地网络的防火墙、代理设置,或者目标服务器所在的网络屏蔽了你的IP,导致无法建立TCP连接。可直接在浏览器中访问附件URL,排查是否能正常下载。
  • Scrapy下载器配置不足:仅设置DOWNLOAD_DELAY不够,需调整DOWNLOAD_TIMEOUT延长超时时间,或降低CONCURRENT_REQUESTS减少并发请求数,避免触发服务器防御机制。
  • 服务器状态异常:目标服务器的文件存储节点出现故障,导致无法响应附件下载请求。
  • URL拼接问题:response.follow可能拼接出错误的附件URL,可打印实际请求的URL确认是否正确。

内容的提问来源于stack exchange,提问作者Mudassir Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 17:10:03