You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas、Matplotlib、Seaborn高效绘制多时序细线条平滑图?

Great question! Let's break down your issues and fix them step by step to get that smooth, thin line visualization you want—even with 1M rows of data.

First, let's tackle the legend problem: when you used size=0.1 in sns.lineplot, Seaborn interpreted that as a variable to map line sizes to, hence the extra numerical entries cluttering your legend. Instead of using size, you should set a fixed thin line width directly with the linewidth parameter (e.g., linewidth=0.3).

Next, the performance and aggregation issues: looping through each consumer's DataFrame is slow because each plot call carries significant overhead. Merging all data into one DataFrame is the right approach, but you missed two critical parameters to avoid aggregation and plot individual consumer lines: units and estimator=None.

Here's the step-by-step solution:

1. Combine all data into a single DataFrame with consumer identifiers

First, add a unique consumer_id to each of your individual consumer DataFrames, then concatenate them into one large DataFrame. This lets Seaborn handle all lines in a single, efficient plot call:

import pandas as pd

# Add a unique ID to each consumer's dataset
for consumer_idx, df in enumerate(energyData):
    df['consumer_id'] = consumer_idx

# Merge all datasets into one
combined_df = pd.concat(energyData, ignore_index=True)

2. Plot with Seaborn to avoid aggregation and improve speed

Use the units parameter to tell Seaborn each consumer is a separate time series, and estimator=None to disable aggregation (which was causing those error lines instead of individual lines). Set a thin line width and use hue to distinguish energy types:

import seaborn as sns
import matplotlib.pyplot as plt

fig, ax = plt.subplots(figsize=(12, 6))

# Plot all lines in one optimized call
sns.lineplot(
    x='t',
    y='demand',
    hue='type',  # Color lines by energy type (electricity/heat)
    units='consumer_id',  # Treat each consumer as a separate time series
    estimator=None,  # Disable aggregation (no mean/error bars)
    linewidth=0.3,  # Thin lines for the sleek, dense look you want
    ax=ax,
    legend='brief'  # Simplify legend to only show energy types
)

# Optional: Clean up plot labels for readability
ax.set_xlabel('Time')
ax.set_ylabel('Demand')
ax.set_title('Consumer Energy Demand Over Time')

plt.show()

3. Extra performance optimizations

  • Sort your data: Ensure combined_df is sorted by t and consumer_id before plotting. This helps Matplotlib render lines more efficiently:
    combined_df = combined_df.sort_values(['t', 'consumer_id'])
    
  • Downsample if possible: If your time resolution is higher than needed (e.g., 1-second intervals when 1-minute is sufficient), downsample the data to reduce the total number of points. Use pandas.resample if t is a datetime, or aggregate by time bins if it's a numerical value.

Why this works

  • The units parameter tells Seaborn to draw a separate line for each unique consumer, while hue keeps electricity and heat lines distinct by color.
  • estimator=None prevents Seaborn from averaging demand across consumers, which was the root cause of those error bars instead of individual lines.
  • Using a fixed linewidth eliminates the spurious size entries in the legend and gives you the thin, clean lines you're aiming for.

This approach should handle your 1M-row dataset efficiently and produce the smooth, dense line visualization you're looking for.

内容的提问来源于stack exchange,提问作者ktnr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 14:18:10