Python从Google Search Console导出数据到BigQuery超时问题求助
解决Google Search Console数据导出到BigQuery的超时问题
在使用Python将Google Search Console(GSC)数据导出到BigQuery时,持续出现超时错误,具体报错信息如下:
Traceback (most recent call last): File "/Users/markleach/Python/GSC_BQ/gsc_bq.py", line 147, in <module> y = get_sc_df(p,"2021-12-01","2022-12-01",x) File "/Users/markleach/Python/GSC_BQ/gsc_bq.py", line 71, in get_sc_df response = service.searchanalytics().query(siteUrl=site_url, body=request).execute() File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/googleapiclient/_helpers.py", line 130, in positional_wrapper return wrapped(*args, **kwargs) File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/googleapiclient/http.py", line 923, in execute resp, content = _retry_request( File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/googleapiclient/http.py", line 222, in _retry_request raise exception File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/googleapiclient/http.py", line 191, in _retry_request resp, content = http.request(uri, method, *args, **kwargs) File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/google_auth_httplib2.py", line 218, in request response, content = self.http.request( File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/httplib2/__init__.py", line 1720, in request (response, content) = self._request( File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/httplib2/__init__.py", line 1440, in _request (response, content) = self._conn_request(conn, request_uri, method, body, headers) File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/site-packages/httplib2/__init__.py", line 1392, in _conn_request response = conn.getresponse() File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/http/client.py", line 1377, in getresponse response.begin() File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/http/client.py", line 320, in begin version, status, reason = self._read_status() File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/http/client.py", line 281, in _read_status line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1") File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/socket.py", line 704, in readinto return self._sock.recv_into(b) File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/ssl.py", line 1242, in recv_into return self.read(nbytes, buffer) File "/Users/markleach/opt/anaconda3/envs/Sandpit/lib/python3.9/ssl.py", line 1100, in read return self._sslobj.read(len, buffer) socket.timeout: The read operation timed out
相关调用代码:
for p in PROPERTIES: for x in range(0,100000,25000): y = get_sc_df(p,"2021-12-01","2022-12-01",x) if len(y) < 25000: break else: continue
解决建议
延长请求超时时间:在API请求的
execute()方法中直接设置超时参数,给服务器足够时间返回数据。修改示例:response = service.searchanalytics().query(siteUrl=site_url, body=request).execute(timeout=60) # 设置60秒超时也可以在构建HTTP客户端时全局配置超时:
import httplib2 http = httplib2.Http(timeout=60) # 后续用该http对象构建Google API服务减小单次请求数据量:当前每次请求25000条数据,可能因数据规模过大导致超时。尝试降低
pageSize参数(比如改为10000),减少单批次返回的数据量,分更多次请求获取完整数据。添加指数退避重试机制:针对
socket.timeout这类临时网络错误,实现自动重试逻辑。可以用tenacity库简化实现:import socket from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type @retry(stop=stop_after_attempt(5), # 最多重试5次 wait=wait_exponential(multiplier=1, min=2, max=10), # 指数退避等待:2s,4s,8s...最大10s retry=retry_if_exception_type(socket.timeout)) def get_sc_df(site_url, start_date, end_date, start_row): # 原函数逻辑不变,仅添加装饰器 request = { 'startDate': start_date, 'endDate': end_date, 'startRow': start_row, 'rowLimit': 25000 # 可根据情况调整 # 其他请求参数 } response = service.searchanalytics().query(siteUrl=site_url, body=request).execute(timeout=60) # 处理response并返回DataFrame # ...拆分时间范围:当前请求的时间跨度为1年,数据量可能远超API处理能力。将时间范围拆分为按月/按周的小批次,比如每月请求一次,降低单次请求的数据规模。
排查网络环境:检查本地网络稳定性,避免因网络波动导致的超时。必要时切换网络或使用稳定的代理服务。
内容的提问来源于stack exchange,提问作者Mark Leach
相关产品推荐
相关产品推荐

