You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

新冠病例数据柱状图定期出现异常尖峰的原因及修正方法

Why Am I Seeing False Spikes in My COVID-19 New Cases Bar Chart?

Let's break down exactly what's causing those incorrect spikes and how to fix your chart step by step.

Root Causes of the False Spikes

  1. Unapplied Missing Value Handling
    Your code calls df.fillna(0) but doesn't assign the result back to df. This means any NaN values in your new_cases column are still present—Matplotlib can render these incorrectly, creating visual artifacts that look like spikes even when the actual value should be 0.

  2. Missing Import for Date Locator
    You're using mdates.WeekdayLocator but haven't imported matplotlib.dates as mdates. While this would normally throw an error, if you somehow ran the code without fixing this, it might fall back to a default locator that misaligns dates, contributing to the spike appearance.

  3. Potential Date Index Gaps
    If your dataset has missing dates (even if new_cases is 0 for those days), the bar chart will draw columns with width proportional to the time between dates. Large gaps can make adjacent bars appear merged or exaggerated, creating false spikes.

Fixed Code to Generate the Correct Chart

Here's the revised code that addresses all these issues:

import pandas as pd
import matplotlib.pyplot as plt
import numpy as np
from matplotlib.dates import DateFormatter
import matplotlib.dates as mdates  # Add missing mdates import

# Load and clean the data correctly
df = pd.read_csv("CovidIndiaData.csv", parse_dates=['date'], index_col=['date'])
df = df[['new_cases', 'total_cases']]
df = df.fillna(0)  # Assign the filled DataFrame back to df to apply changes!

# Ensure continuous daily date range (fill missing dates with 0 new cases)
full_date_range = pd.date_range(start='2020-01-01', end='2020-07-18', freq='D')
df = df.reindex(full_date_range, fill_value=0)

# Create the chart
fig = plt.figure(figsize=(12, 6))  # Larger figure for better readability
ax = plt.gca()

ax.bar(df.index.values, df['new_cases'], color='purple', width=0.8)  # Explicit width to avoid overlap
ax.set(xlabel="Date", ylabel="New Cases", title="New Cases per Day (Jan-Jul 2020)", 
       xlim=["2020-01-01", "2020-07-18"])

# Format date ticks properly
date_form = DateFormatter("%m-%d")
ax.xaxis.set_major_formatter(date_form)
ax.xaxis.set_major_locator(mdates.WeekdayLocator(interval=1))  # Now works with proper import

plt.setp(ax.get_xticklabels(), rotation=45, ha='right')  # Prevent label cutoff
plt.tight_layout()  # Adjust layout to fit rotated labels
plt.show()

Key Fixes Explained

  • Applied fillna(0): By assigning df = df.fillna(0), we ensure all missing values are replaced with 0, eliminating unexpected rendering from NaNs.
  • Continuous Date Index: Using reindex with a full daily date range fills in any missing dates with 0 new cases, so bars are drawn consistently for every day without gaps that cause visual spikes.
  • Proper Date Locator Import: We added the missing mdates import so the weekday locator works as intended, keeping date ticks aligned correctly.
  • Improved Label Readability: Added ha='right' to rotated tick labels and tight_layout() to prevent cutoff, making the chart easier to interpret.

Verification Tip

To confirm the fix, check the values for the problematic dates (like 2020-01-07) with this line:

print(df.loc['2020-01-07', 'new_cases'])

This should return 0, matching your dataset's actual values, and the bar for that date will now be flat instead of showing a spike.


内容的提问来源于stack exchange,提问作者blahblahblah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 14:17:46