You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解码Apache Parquet?API获取Parquet数据处理报错求助

解决Parquet格式API数据读取报错问题

问题根源

你用response.text获取Parquet数据是错误的——Parquet是二进制格式,response.text会把原始二进制字节按默认字符编码(比如UTF-8)解析成字符串,这个过程会破坏Parquet的二进制结构,导致后续无法正常解析。

修正方案

1. 修正数据获取逻辑

把获取数据的代码改成直接读取response.content(二进制字节数据),不要转成字符串:

import requests

def get_orders_data(date):
    # 注意:补全完整API域名,比如"https://your-api-domain.com/orders"
    url = "/orders"
    headers = {}
    params = {
        "date": date
    }

    response = requests.get(url, headers=headers, params=params)

    if response.status_code == 200:
        # 直接返回二进制字节数据,跳过字符串转换
        return response.content
    else:
        print("Failed to fetch orders data")
        return None

date_to_fetch = "2023-12-08"
orders_info = get_orders_data(date_to_fetch)

2. 修正Parquet解析逻辑

现在orders_info是原始二进制数据,直接传入BytesIO即可解析,不需要再做encode操作:

import io
import pyarrow.parquet as pq
import pandas as pd

if orders_info:
    # 用二进制数据直接创建内存缓冲区
    buffer = io.BytesIO(orders_info)
    
    # 读取并解析Parquet数据
    table = pq.read_table(buffer)
    df = table.to_pandas()
    
    # 验证数据
    print(df.head())

    # 可选:保存为本地Parquet文件
    df.to_parquet("orders_20231208.parquet")

扩展:上传到Google Cloud Storage(GCS)

如果需要将文件上传到GCP的GCS,可添加以下代码(需先安装google-cloud-storage库):

from google.cloud import storage

def upload_to_gcs(bucket_name, source_file_name, destination_blob_name):
    storage_client = storage.Client()
    bucket = storage_client.bucket(bucket_name)
    blob = bucket.blob(destination_blob_name)

    blob.upload_from_filename(source_file_name)
    print(f"File {source_file_name} uploaded to {destination_blob_name}.")

# 调用示例
upload_to_gcs("your-gcs-bucket", "orders_20231208.parquet", "daily_orders/orders_20231208.parquet")

内容的提问来源于stack exchange,提问作者Artem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 08:57:33