如何解码Apache Parquet?API获取Parquet数据处理报错求助
解决Parquet格式API数据读取报错问题
问题根源
你用response.text获取Parquet数据是错误的——Parquet是二进制格式,response.text会把原始二进制字节按默认字符编码(比如UTF-8)解析成字符串,这个过程会破坏Parquet的二进制结构,导致后续无法正常解析。
修正方案
1. 修正数据获取逻辑
把获取数据的代码改成直接读取response.content(二进制字节数据),不要转成字符串:
import requests def get_orders_data(date): # 注意:补全完整API域名,比如"https://your-api-domain.com/orders" url = "/orders" headers = {} params = { "date": date } response = requests.get(url, headers=headers, params=params) if response.status_code == 200: # 直接返回二进制字节数据,跳过字符串转换 return response.content else: print("Failed to fetch orders data") return None date_to_fetch = "2023-12-08" orders_info = get_orders_data(date_to_fetch)
2. 修正Parquet解析逻辑
现在orders_info是原始二进制数据,直接传入BytesIO即可解析,不需要再做encode操作:
import io import pyarrow.parquet as pq import pandas as pd if orders_info: # 用二进制数据直接创建内存缓冲区 buffer = io.BytesIO(orders_info) # 读取并解析Parquet数据 table = pq.read_table(buffer) df = table.to_pandas() # 验证数据 print(df.head()) # 可选:保存为本地Parquet文件 df.to_parquet("orders_20231208.parquet")
扩展:上传到Google Cloud Storage(GCS)
如果需要将文件上传到GCP的GCS,可添加以下代码(需先安装google-cloud-storage库):
from google.cloud import storage def upload_to_gcs(bucket_name, source_file_name, destination_blob_name): storage_client = storage.Client() bucket = storage_client.bucket(bucket_name) blob = bucket.blob(destination_blob_name) blob.upload_from_filename(source_file_name) print(f"File {source_file_name} uploaded to {destination_blob_name}.") # 调用示例 upload_to_gcs("your-gcs-bucket", "orders_20231208.parquet", "daily_orders/orders_20231208.parquet")
内容的提问来源于stack exchange,提问作者Artem
相关产品推荐
相关产品推荐

