You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python花粉数据分析:X轴日期排序问题及日期对应花粉概率查询功能咨询

Python花粉数据分析:X轴日期排序问题及日期对应花粉概率查询功能咨询

Hey Lola!完全理解作为医学生刚接触编程的心情,刚学半学期就着手做数据分析项目已经超棒了,别担心基础问题~我来帮你解决这两个核心问题:

一、解决X轴日期排序混乱的问题

你现在的X轴显示的是(月,日)元组,因为groupby([month, day])默认是按分组的原始顺序排列,没有遵循自然日期逻辑,所以才会出现乱序的情况。我们可以把索引转换成真正的日期对象,再排序,就能让X轴按时间顺序排列了:

具体实现步骤:

  1. 将分组后的元组索引转换为datetime格式
    给每个(月,日)组合加一个虚拟年份(比如2020,闰年能覆盖2月29日的特殊情况),把它转换成标准日期:
    # 先把索引拆成独立的月份、日期列
    df3 = df3.reset_index()
    df3.columns = ['month', 'day'] + df3.columns[2:].tolist()
    
    # 创建虚拟日期列,统一用2020年
    df3['date'] = pd.to_datetime(df3[['month', 'day']].assign(year=2020))
    
    # 按日期排序,再把date设为索引
    df3 = df3.sort_values('date').set_index('date')
    
  2. 重新绘图并优化X轴显示
    现在索引是有序的日期了,绘图时X轴会自动从1月1日到12月31日自然排列。如果想让X轴显示更友好(比如"MM-DD"格式),可以用matplotlib的日期格式化工具:
    import matplotlib.dates as mdates
    
    plt.figure(figsize=(12,6))
    df3['某花粉类型'].plot()  # 替换成你要展示的花粉列名
    
    # 设置X轴显示格式为「月-日」
    plt.gca().xaxis.set_major_formatter(mdates.DateFormatter('%m-%d'))
    # 每隔30天显示一个刻度,避免X轴拥挤
    plt.gca().xaxis.set_major_locator(mdates.DayLocator(interval=30))
    plt.xticks(rotation=45)
    plt.show()
    

这样你的图表X轴就会整齐按时间顺序排列啦~

二、实现「给定日期查询花粉概率」的功能

这里的"概率"可以从两个角度实现,结合你现有的数据给你两种方案:

方案1:基于日均浓度判断

当某类花粉的26年日均浓度超过设定阈值时,我们认为它"很可能存在于空气中":

def get_pollen_by_date(target_date_str):
    # 把输入的日期字符串(比如"2024-05-15")转换成datetime对象
    target_date = pd.to_datetime(target_date_str)
    target_month = target_date.month
    target_day = target_date.day

    # 构造虚拟日期(和之前的年份一致,确保能匹配df3的索引)
    dummy_date = pd.to_datetime(f'2020-{target_month}-{target_day}')
    daily_mean = df3.loc[dummy_date]

    # 设定阈值:可以用所有花粉日均浓度的中位数,也可以自定义(比如>0就认为存在)
    threshold = daily_mean.median()

    # 筛选出超过阈值的花粉
    likely_pollens = daily_mean[daily_mean > threshold].index.tolist()

    if likely_pollens:
        print(f"在{target_date_str},以下花粉很可能存在于空气中:")
        for pollen in likely_pollens:
            print(f"- {pollen}: 26年日均浓度 {daily_mean[pollen]:.2f}")
    else:
        print(f"在{target_date_str},没有检测到高概率存在的花粉")

使用示例:

get_pollen_by_date("2024-05-10")

方案2:基于年份占比计算真实概率

统计26年中该日期有多少天检测到了花粉,计算占比作为"存在概率",更贴近你说的"概率"定义:

def get_pollen_probability(target_date_str):
    target_date = pd.to_datetime(target_date_str)
    target_month = target_date.month
    target_day = target_date.day

    # 统计每年该日期的花粉是否存在(浓度>0即视为存在)
    df2['is_present'] = df2[df2.columns[1:]] > 0  # 假设第1列是日期,后面是花粉列
    presence_count = df2.groupby([df2.Date.dt.month, df2.Date.dt.day])['is_present'].sum()

    # 计算概率(存在天数/总年份数26)
    prob_df = presence_count.loc[(target_month, target_day)] / 26
    # 筛选概率超过50%的花粉
    high_prob_pollens = prob_df[prob_df > 0.5].index.tolist()

    if high_prob_pollens:
        print(f"在{target_date_str},以下花粉存在概率较高:")
        for pollen in high_prob_pollens:
            print(f"- {pollen}: {prob_df[pollen]:.1%}")
    else:
        print(f"在{target_date_str},没有存在概率超过50%的花粉")

希望这些方法能帮到你,加油!

备注:内容来源于stack exchange,提问作者Lola Scheuren

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.22 11:52:57