You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python从Google Search Console导出数据到BigQuery超时问题求助

解决Google Search Console数据导出到BigQuery的超时问题

在使用Python将Google Search Console(GSC)数据导出到BigQuery时,持续出现超时错误,具体报错信息如下:

Traceback (most recent call last):
  File "/Users/markleach/Python/GSC_BQ/gsc_bq.py", line 147, in <module>
    y = get_sc_df(p,"2021-12-01","2022-12-01",x)
  File "/Users/markleach/Python/GSC_BQ/gsc_bq.py", line 71, in get_sc_df
    response = service.searchanalytics().query(siteUrl=site_url, body=request).execute()
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/googleapiclient/_helpers.py", line 130, in positional_wrapper
    return wrapped(*args, **kwargs)
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/googleapiclient/http.py", line 923, in execute
    resp, content = _retry_request(
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/googleapiclient/http.py", line 222, in _retry_request
    raise exception
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/googleapiclient/http.py", line 191, in _retry_request
    resp, content = http.request(uri, method, *args, **kwargs)
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/google_auth_httplib2.py", line 218, in request
    response, content = self.http.request(
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/httplib2/__init__.py", line 1720, in request
    (response, content) = self._request(
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/httplib2/__init__.py", line 1440, in _request
    (response, content) = self._conn_request(conn, request_uri, method, body, headers)
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/httplib2/__init__.py", line 1392, in _conn_request
    response = conn.getresponse()
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/http/client.py", line 1377, in getresponse
    response.begin()
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/http/client.py", line 320, in begin
    version, status, reason = self._read_status()
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/http/client.py", line 281, in _read_status
    line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1")
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/socket.py", line 704, in readinto
    return self._sock.recv_into(b)
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/ssl.py", line 1242, in recv_into
    return self.read(nbytes, buffer)
  File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/ssl.py", line 1100, in read
    return self._sslobj.read(len, buffer)
socket.timeout: The read operation timed out

相关调用代码:

for p in PROPERTIES:

    for x in range(0,100000,25000):
        y = get_sc_df(p,"2021-12-01","2022-12-01",x)
        if len(y) < 25000:
            break
        else:
            continue

解决建议

  • 延长请求超时时间:在API请求的execute()方法中直接设置超时参数,给服务器足够时间返回数据。修改示例:

    response = service.searchanalytics().query(siteUrl=site_url, body=request).execute(timeout=60)  # 设置60秒超时
    

    也可以在构建HTTP客户端时全局配置超时:

    import httplib2
    http = httplib2.Http(timeout=60)
    # 后续用该http对象构建Google API服务
    
  • 减小单次请求数据量:当前每次请求25000条数据,可能因数据规模过大导致超时。尝试降低pageSize参数(比如改为10000),减少单批次返回的数据量,分更多次请求获取完整数据。

  • 添加指数退避重试机制:针对socket.timeout这类临时网络错误,实现自动重试逻辑。可以用tenacity库简化实现:

    import socket
    from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type
    
    @retry(stop=stop_after_attempt(5),  # 最多重试5次
           wait=wait_exponential(multiplier=1, min=2, max=10),  # 指数退避等待:2s,4s,8s...最大10s
           retry=retry_if_exception_type(socket.timeout))
    def get_sc_df(site_url, start_date, end_date, start_row):
        # 原函数逻辑不变,仅添加装饰器
        request = {
            'startDate': start_date,
            'endDate': end_date,
            'startRow': start_row,
            'rowLimit': 25000  # 可根据情况调整
            # 其他请求参数
        }
        response = service.searchanalytics().query(siteUrl=site_url, body=request).execute(timeout=60)
        # 处理response并返回DataFrame
        # ...
    
  • 拆分时间范围:当前请求的时间跨度为1年,数据量可能远超API处理能力。将时间范围拆分为按月/按周的小批次,比如每月请求一次,降低单次请求的数据规模。

  • 排查网络环境:检查本地网络稳定性,避免因网络波动导致的超时。必要时切换网络或使用稳定的代理服务。

内容的提问来源于stack exchange,提问作者Mark Leach

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 12:20:26