AWS S3上传音频分片时遭遇TimeoutError的问题排查与性能优化求助
AWS S3上传音频分片时遭遇TimeoutError的问题排查与性能优化求助
我正在为AWS OpenData项目在S3上构建数据集,流程是先从S3拉取已存储的音频文件,切割成更小的音频分片后,再上传回S3。但只要执行PUT操作就会出问题——如果注释掉PUT相关代码,就不会报错。
我的S3桶位于美国区域,为了排查是不是我这边的网络问题,我在三个环境做了测试:
- 美国区域的SageMaker免费实验室:没有报错,但程序挂起4小时完全没有进展
- 美国区域的Google Colab:遇到同样的问题,还因为数据量过大被临时限制了资源,现在没法再测试
- 欧盟区域的本地环境:直接返回TimeoutError,完全无法推进
我现在需要解决这个超时错误,同时尽可能提升操作速度。而且受预算限制,只能使用基础S3服务,不能用Lambda等其他AWS服务。
下面是我目前的实现代码和尝试过的方案:
初始化S3客户端
# boto3 client for uploading (signed requests) s3_client = boto3.client( 's3', aws_access_key_id=AWS_ACCESS_KEY, aws_secret_access_key=AWS_SECRET_KEY, region_name=REGION_NAME )
核心处理代码
%%time import functools cache_audio = {} for o, (lb, up) in enumerate(batches[6:]): for ix, row in annotated_segments.loc[lb:up].iterrows(): # 清理缓存以保证内存安全 if len(cache_audio) > 5: cache_audio.clear() # 音频属性 file_name = row['File name'] # 路径属性 file_folder = row['File folder'] # 分片属性 segment_name = row['segment_name'] start = row['voice_start'] end = row['voice_end'] # 从缓存读取音频 if file_name not in cache_audio: audio, rate = fetch_audio(row) cache_audio[file_name] = audio else: audio = cache_audio[file_name] # 提取音频分片 audio_segment = audio[start : end] try: s3_path = f"data/annotated_segments/{file_folder}/{file_name}/{segment_name}" # 初始化二进制文件对象 file_obj = io.BytesIO() # 写入音频分片 soundfile.write(file_obj, audio_segment, samplerate = rate, format='WAV') # 将文件指针重置到开头 file_obj.seek(0) # 上传分片到S3 put_audio_to_s3(file_obj, s3_path) except Exception as e: print(f"上传文件出错: {e}。文件名: { file_name }。批次: {lb} - {up}") print(f"成功完成第{o}批次: {lb} - {up}")
遇到的错误信息
---------------------------------------------------------------------------TimeoutError Traceback (most recent call last)File ~/miniconda3/envs/fruitbats/lib/python3.10/site-packages/urllib3/response.py:754, in HTTPResponse._error_catcher(self) 753 try:--> 754 yield 756 except SocketTimeout as e: 757 # FIXME: Ideally we'd like to include the url in the ReadTimeoutError but 758 # there is yet no clean way to get at it from this context.File ~/miniconda3/envs/fruitbats/lib/python3.10/site-packages/urllib3/response.py:879, in HTTPResponse._raw_read(self, amt, read1) 878 with self._error_catcher():--> 879 data = self._fp_read(amt, read1=read1) if not fp_closed else b"" 880 if amt is not None and amt != 0 and not data: 881 # Platform-specific: Buggy versions of Python. 882 # Close the connection when no data is returned (...) 887 # not properly close the connection in all cases. There is 888 # no harm in redundantly calling close.File ~/miniconda3/envs/fruitbats/lib/python3.10/site-packages/urllib3/response.py:862, in HTTPResponse._fp_read(self, amt, read1) 860 else: 861 # StringIO doesn't like amt=None--> 862 return self._fp.read(amt) if amt is not None else self._fp.read()File ~/miniconda3/envs/fruitbats/lib/python3.10/http/client.py:482, in HTTPResponse.read(self, amt) 481 try:--> 482 s = self._safe_read(self.length) 483 except IncompleteRead:File ~/miniconda3/envs/fruitbats/lib/python3.10/http/client.py:631, in HTTPResponse._safe_read(self, amt) 625 """Read the number of bytes requested. 626 627 This function should be used when <amt> bytes "should" be present for 628 reading. If the bytes are truly not available (due to EOF), then the 629 IncompleteRead exception can be used to detect the problem. 630 """--> 631 data = self.fp.read(amt) 632 if len(data) < amt:File ~/miniconda3/envs/fruitbats/lib/python3.10/socket.py:717, in SocketIO.readinto(self, b) 716 try:--> 717 return self._sock.recv_into(b) 718 except timeout:File ~/miniconda3/envs/fruitbats/lib/python3.10/ssl.py:1307, in SSLSocket.recv_into(self, buffer, nbytes, flags) 1304 raise ValueError( 1305 "non-zero flags not allowed in calls to recv_into() on %s" % 1306 self.__class__)--> 1307 return self.read(nbytes, buffer) 1308 else:File ~/miniconda3/envs/fruitbats/lib/python3.10/ssl.py:1163, in SSLSocket.read(self, len, buffer) 1162 if buffer is not None:--> 1163 return self._sslobj.read(len, buffer) 1164 else:TimeoutError: The read operation timed outThe above exception was the direct cause of the following exception:ReadTimeoutError Traceback (most recent call last)File ~/miniconda3/envs/fruitbats/lib/python3.10/site-packages/botocore/response.py:99, in StreamingBody.read(self, amt) 98 try:---> 99 chunk = self._raw_stream.read(amt) 100 except URLLib3ReadTimeoutError as e: 101 # TODO: the url will be None as urllib3 isn't setting it yetFile ~/miniconda3/envs/fruitbats/lib/python3.10/site-packages/urllib3/response.py:955, in HTTPResponse.read(self, amt, decode_content, cache_content) 953 return self._decoded_buffer.get(amt)--> 955 data = self._raw_read(amt) 957 flush_decoder = amt is None or (amt != 0 and not data)File ~/miniconda3/envs/fruitbats/lib/python3.10/site-packages/urllib3/response.py:878, in HTTPResponse._raw_read(self, amt, read1) 876 fp_closed = getattr(self._fp, "closed", False)--> 878 with self._error_catcher(): 879 data = self._fp_read(amt, read1=read1) if not fp_closed else b""File ~/miniconda3/envs/fruitbats/lib/python3.10/contextlib.py:153, in _GeneratorContextManager.__exit__(self, typ, value, traceback) 152 try:--> 153 self.gen.throw(typ, value, traceback) 154 except StopIteration as exc: 155 # Suppress StopIteration *unless* it's the same exception that 156 # was passed to throw(). This prevents a StopIteration 157 # raised inside the "with" statement from being suppressed.File ~/miniconda3/envs/fruitbats/lib/python3.10/site-packages/urllib3/response.py:759, in HTTPResponse._error_catcher(self) 756 except SocketTimeout as e: 757 # FIXME: Ideally we'd like to include the url in the ReadTimeoutError but 758 # there is yet no clean way to get at it from this context.--> 759 raise ReadTimeoutError(self._pool, None, "Read timed out.") from e # type: ignore[arg-type] 761 except BaseSSLError as e: 762 # FIXME: Is there a better way to differentiate between SSLErrors?ReadTimeoutError: AWSHTTPSConnectionPool(host='fruitbat-vocalizations.s3.us-west-2.amazonaws.com', port=443): Read timed out.During handling of the above exception, another exception occurred:ReadTimeoutError Traceback (most recent call last)File <timed exec>:25Cell In[25], line 15, in fetch_audio(row, sr) 12 s3_object_key = str(s3_path.relative_to(DSLOC)) 14 response = s3_client.get_object(Bucket=BUCKET_NAME, Key=s3_object_key)---> 15 file_content = response['Body'].read() 17 # https://stackoverflow.com/questions/73350508/read-audio-file-from-s3-directly-in-python 18 # this will read in float64 by default and multichannel if any 19 data, rate = soundfile.read(io.BufferedReader(io.BytesIO(file_content)), always_2d=True)File ~/miniconda3/envs/fruitbats/lib/python3.10/site-packages/botocore/httpchecksum.py:240, in StreamingChecksumBody.read(self, amt) 239 def read(self, amt=None):--> 240 chunk = super().read(amt=amt) 241 self._checksum.update(chunk) 242 if amt is None or (not chunk and amt > 0):File ~/miniconda3/envs/fruitbats/lib/python3.10/site-packages/botocore/response.py:102, in StreamingBody.read(self, amt) 99 chunk = self._raw_stream.read(amt) 100 except URLLib3ReadTimeoutError as e: 101 # TODO: the url will be None as urllib3 isn't setting it yet--> 102 raise ReadTimeoutError(endpoint_url=e.url, error=e) 103 except URLLib3ProtocolError as e: 104 raise ResponseStreamingError(error=e)ReadTimeoutError: Read timeout on endpoint URL: "None"
希望社区的大佬们能帮忙解决这个超时问题,同时给点提升操作效率的建议,万分感谢!
备注:内容来源于stack exchange,提问作者user305883
相关产品推荐
相关产品推荐

