You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Airflow XCom序列化报错:DataFrame无法转为JSON可序列化对象的原因解析与解决方案

Airflow XCom Error: DataFrame is not JSON serializable

Let's break down your questions and fix this issue step by step:

1. What does "JSON serializable" mean?

JSON (JavaScript Object Notation) is a lightweight data exchange format that only supports a limited set of basic types: strings, numbers, booleans, arrays, key-value objects, and null.

JSON serializable refers to an object's ability to be converted into a valid JSON string. This conversion lets data be stored, transferred, and later reconstructed back into its original structure. If an object’s type isn’t natively supported by JSON and there’s no built-in conversion logic for it, you’ll hit the "not JSON serializable" error.

2. How does your DataFrame relate to JSON here?

In Airflow, when a PythonOperator task returns a value, that value is automatically stored by XCom (Cross-Communication)—Airflow’s built-in system for sharing data between tasks. By default, XCom uses JSON to serialize and store these values.

A Pandas DataFrame is a complex Python object: it holds not just your raw data, but also metadata like indexes, column data types, and formatting rules. JSON has no native way to represent this layered structure, so the JSON encoder can’t convert the DataFrame into a valid JSON string—hence the error you’re seeing.

3. How to fix this error?

You have a few practical options depending on your use case:

Option 1: Convert the DataFrame to a JSON string

Turn your DataFrame into a JSON string before returning it, then parse it back into a DataFrame in your downstream task. This works with Airflow’s default JSON-based XCom setup.

Update your code like this:

def _game():
    game_df = pd.DataFrame(columns=['match', 'puuid'])
    input_data = {'match': match100, 'puuid': game_content['puuid']}
    game_df = game_df.append(input_data, ignore_index=True)
    # Convert DataFrame to a row-based JSON string
    return game_df.to_json(orient='records')

def _game_score(**context):
    # Pull the JSON string from XCom and convert back to DataFrame
    game_json = context['ti'].xcom_pull(task_ids='create_data_df')
    game_df = pd.read_json(game_json)
    # Add your data processing logic here...

The orient='records' parameter formats the JSON as an array of row objects, which is ideal for your tabular data.

Option 2: Enable Pickle serialization for XCom

If you want to pass the DataFrame directly without conversion, you can switch XCom to use Pickle—a Python-specific serialization format that can handle almost any Python object. Important Note: Pickle is insecure for untrusted data, as malicious Pickle files can execute arbitrary code. Only use this in environments you fully control.

Edit your airflow.cfg configuration file:

[core]
enable_xcom_pickling = True

Restart Airflow after making this change. Now your tasks can return DataFrames directly, and downstream tasks will receive the full object via xcom_pull:

def _game_score(**context):
    game_df = context['ti'].xcom_pull(task_ids='create_data_df')
    # Use the DataFrame directly for processing...

Option 3: Use external storage (best for large datasets)

XCom isn’t designed for large datasets—it stores data in Airflow’s metadata database, which can slow down with big DataFrames. Instead, save your DataFrame to an external store (like a CSV file, cloud storage bucket, or database) and pass the file/path reference via XCom.

Example with a local CSV file (adjust paths for your environment):

def _game():
    game_df = pd.DataFrame(columns=['match', 'puuid'])
    input_data = {'match': match100, 'puuid': game_content['puuid']}
    game_df = game_df.append(input_data, ignore_index=True)
    # Save DataFrame to a CSV file
    file_path = '/opt/airflow/data/game_data.csv'
    game_df.to_csv(file_path, index=False)
    # Return the file path for downstream tasks to use
    return file_path

def _game_score(**context):
    file_path = context['ti'].xcom_pull(task_ids='create_data_df')
    game_df = pd.read_csv(file_path)
    # Add your data processing logic here...

内容的提问来源于stack exchange,提问作者lima

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 14:49:07