Python3如何从URL读取gzip文件并使用islice逐行提取数据
需求概述
需要读取托管在网站上、压缩后大小超过20GB的txt.gz文件,希望直接用gzip打开后调用itertools的islice逐行提取内容,目前gzip原生不支持直接读取远程URL文件。
遇到的问题
urllib这类库默认会一次性下载完整二进制数据流,目前找到的使用urllib或requests的脚本都是先将流下载到本地文件或变量后再解压读取文本。由于数据集过大,需要边下载边处理;同时需要按行迭代文本,按字节设置下载块大小无法保证数据边界为换行符(数据固定为换行符分隔)。
本地可用示例代码(不支持URL读取)
如下代码在处理本地磁盘文件时运行效果很好:
from itertools import islice import gzip # Gzip file open call datafile = gzip.open("/home/shrout/Documents/line_numbers.txt.gz") chunk_size = 2 while True: data_chunk = list(islice(datafile, chunk_size)) if not data_chunk: break print(data_chunk) datafile.close()
脚本示例输出
shrout@ubuntu:~/Documents$ python3 itertools_test.py [b'line 1\n', b'line 2\n'] [b'line 3\n', b'line 4\n'] [b'line 5\n', b'line 6\n'] [b'line 7\n', b'line 8\n'] [b'line 9\n', b'line 10\n'] [b'line 11\n', b'line 12\n'] [b'line 13\n', b'line 14\n'] [b'line 15\n', b'line 16\n'] [b'line 17\n', b'line 18\n'] [b'line 19\n', b'line 20\n']
适配需求说明
现有相关方案都没有在处理流的过程中解压读取数据,都是将二进制数据直接写入新的本地文件或脚本变量。该数据集过大无法全部载入内存,提前将原文件写入磁盘再读取也会浪费时间。
目前可以在虚拟机上用上述示例代码本地完成任务,但现在需要适配minio对象存储和docker容器环境,需要找到可以被gzip.open或同类方法直接调用的URL文件句柄实现方案。
已实现的部分方案
目前已经实现了分块流式下载gzip文件并解压的代码,不过将数据拆分为换行分隔的字符串还需要额外的处理开销,目前还没有找到最优方案:
import requests import zlib target_url = "http://127.0.0.1:9000/test-bucket/big_data_file.json.gz" # Using zlib.MAX_WBITS|32 apparently forces zlib to detect the appropriate header for the data decompressor = zlib.decompressobj(zlib.MAX_WBITS|32) # Stream this file in as a request - pull the content in just a little at a time with requests.get(target_url, stream=True) as remote_file: # Chunk size can be adjusted to test performance for chunk in remote_file.iter_content(chunk_size=8192): # Decompress the current chunk decompressed_chunk = decompressor.decompress(chunk) print(decompressed_chunk)
内容的提问来源于stack exchange,提问作者Shrout1
相关产品推荐
相关产品推荐

