You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Flask环境中实现跨会话共享可编辑的Pandas DataFrame?

How to Share a Single Pandas DataFrame Across All Flask API Clients

Great question! The key here is to create a globally accessible, thread-safe DataFrame that all clients interact with, without relying on per-session storage or problematic global variables. Let's break down the solutions for both development and production environments.

1. Development Environment (Single Process)

For local testing or single-process deployments, you can attach the DataFrame and a thread lock directly to your Flask app instance. This avoids the issues with standard global variables (which can behave unpredictably in Flask's request context) and ensures all requests modify the same data.

Step-by-Step Implementation

from flask import Flask
import pandas as pd
import threading

app = Flask(__name__)

# Initialize shared DataFrame and thread lock when the app starts
def init_shared_data():
    # Load your JSON file into a DataFrame
    app.shared_df = pd.read_json('your_data.json')
    # Add a lock to prevent race conditions from concurrent requests
    app.df_lock = threading.Lock()

# Run initialization once when the app starts
init_shared_data()

# Example route: Delete a column (visible to all clients)
@app.route('/delete-column/<col_name>', methods=['DELETE'])
def delete_column(col_name):
    # Use the lock to ensure only one request modifies the DataFrame at a time
    with app.df_lock:
        if col_name in app.shared_df.columns:
            app.shared_df.drop(col_name, axis=1, inplace=True)
            return {"message": f"Column '{col_name}' deleted successfully"}, 200
        return {"error": "Column not found"}, 404

# Example route: Get the current state of the DataFrame
@app.route('/get-data', methods=['GET'])
def get_data():
    with app.df_lock:
        # Convert DataFrame to JSON for the response
        return app.shared_df.to_json(orient='records'), 200

Why This Works

  • The DataFrame is stored as an attribute of the Flask app instance, which is a single object shared across all requests in a single process.
  • The threading.Lock() prevents race conditions (e.g., two clients trying to modify the DataFrame at the same time, leading to corrupted data).

2. Production Environment (Multi-Process/Server)

If you deploy your Flask app with multiple workers (e.g., Gunicorn with --workers 4), each worker runs in its own process with its own copy of the app instance. This means changes to the DataFrame in one worker won't be visible to others. For this scenario, you need an external shared storage system.

Option A: Redis (In-Memory Data Store)

Redis is perfect for storing lightweight, frequently accessed data like your DataFrame. It supports atomic operations and distributed locking to handle concurrency.

from flask import Flask
import pandas as pd
import redis
from redis.lock import Lock

app = Flask(__name__)
# Connect to your Redis instance (local or remote)
redis_client = redis.Redis(host='localhost', port=6379, db=0)

# Initialize: Load JSON data into Redis once
def init_redis_data():
    df = pd.read_json('your_data.json')
    # Store DataFrame as JSON in Redis
    redis_client.set('shared_df', df.to_json(orient='records'))

init_redis_data()

@app.route('/delete-column/<col_name>', methods=['DELETE'])
def delete_column(col_name):
    # Use a distributed lock to ensure atomicity across processes
    with Lock(redis_client, 'df_lock'):
        # Fetch current data from Redis
        df_json = redis_client.get('shared_df')
        if not df_json:
            return {"error": "Data not found"}, 500
        
        df = pd.read_json(df_json, orient='records')
        if col_name in df.columns:
            df.drop(col_name, axis=1, inplace=True)
            # Save updated DataFrame back to Redis
            redis_client.set('shared_df', df.to_json(orient='records'))
            return {"message": f"Column '{col_name}' deleted successfully"}, 200
        return {"error": "Column not found"}, 404

@app.route('/get-data', methods=['GET'])
def get_data():
    df_json = redis_client.get('shared_df')
    if not df_json:
        return {"error": "Data not found"}, 500
    return df_json, 200, {'Content-Type': 'application/json'}

Option B: Relational Database (e.g., PostgreSQL)

If your data is complex or needs persistent storage (even after app restarts), store the DataFrame's data in a database table. All clients will interact with the same table, ensuring consistency.

  • Use df.to_sql() to load initial data into the database.
  • For modifications, execute SQL queries (e.g., ALTER TABLE ... DROP COLUMN) or use Pandas to read, modify, and write back (with transactions for safety).

Why Session Storage Isn't Right for This

Session storage is designed to isolate data per client, so changes made by one client won't affect others. That's exactly what you're trying to avoid here—so it's not a fit for shared state across all clients.

内容的提问来源于stack exchange,提问作者Tochi Bedford

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 14:37:29