You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在存在缺失值时用线条连接Biomarker数据点

问题描述

需要在同一张图中绘制多个Biomarker随日期的变化曲线,但各Biomarker的采样日期不同且存在缺失值。当前绘图代码因缺失值导致线条断裂,要求直接连接同一Biomarker的所有有效数据点,且禁止使用interpolate生成新数据点。

示例数据:

data = {
    'PatientID': [244651, 244651, 244651, 244651, 244652, 244653, 244651],
    'LocationType': ['IP', 'IP', 'OP', 'IP', 'IP', 'OP', 'IP'],
    'Date': ['2023-01-01', '2023-01-02', '2023-01-03', '2023-01-04', '2023-01-01', '2023-01-01', '2023-01-05'],
    'Biomarker1': [1.1, 1.2, None, 1.4, 2.1, 3.1, 1.5],
    'Biomarker2': [2.1, None, 2.3, 2.4, 3.1, 4.1, 2.5],
    'Biomarker3': [3.1, 3.2, 3.3, None, 4.1, 5.1, 3.5]
}

原完整代码:

import pandas as pd
import matplotlib.pyplot as plt
import numpy as np

# Sample DataFrame (replace this with your actual DataFrame)
data = {
    'PatientID': [244651, 244651, 244651, 244651, 244652, 244653, 244651],
    'LocationType': ['IP', 'IP', 'OP', 'IP', 'IP', 'OP', 'IP'],
    'Date': ['2023-01-01', '2023-01-02', '2023-01-03', '2023-01-04', '2023-01-01', '2023-01-01', '2023-01-05'],
    'Biomarker1': [1.1, 1.2, None, 1.4, 2.1, 3.1, 1.5],
    'Biomarker2': [2.1, None, 2.3, 2.4, 3.1, 4.1, 2.5],
    'Biomarker3': [3.1, 3.2, 3.3, None, 4.1, 5.1, 3.5]
}

# Create DataFrame
df = pd.DataFrame(data)
df['Date'] = pd.to_datetime(df['Date'])

# Filter the data for the specified patient ID and IP location type
filtered_df = df[(df['PatientID'] == 244651) & (df['LocationType'] == 'IP')]

# Set the date as the index
filtered_df.set_index('Date', inplace=True)

# Plot all biomarkers
plt.figure(figsize=(12, 8))

# Loop through each biomarker column to plot each one separately
for column in filtered_df.columns:
    if column not in ['PatientID', 'LocationType']:
        plt.plot(filtered_df.index, filtered_df[column], marker='o', linestyle='-', label=column)

plt.title('Biomarkers by Date for Patient ID 244651 (IP Location Type)')
plt.xlabel('Date')
plt.ylabel('Biomarker Value')
plt.legend()
plt.grid(True)
plt.xticks(rotation=45)
plt.show()
解决方案

核心方法是对每个Biomarker列单独筛选出无缺失值的行,仅保留有效数据点后再绘图,这样matplotlib会直接连接这些有效点,不会因缺失值断裂,也不会生成新数据。

修改后的完整代码:

import pandas as pd
import matplotlib.pyplot as plt
import numpy as np

# Sample DataFrame
data = {
    'PatientID': [244651, 244651, 244651, 244651, 244652, 244653, 244651],
    'LocationType': ['IP', 'IP', 'OP', 'IP', 'IP', 'OP', 'IP'],
    'Date': ['2023-01-01', '2023-01-02', '2023-01-03', '2023-01-04', '2023-01-01', '2023-01-01', '2023-01-05'],
    'Biomarker1': [1.1, 1.2, None, 1.4, 2.1, 3.1, 1.5],
    'Biomarker2': [2.1, None, 2.3, 2.4, 3.1, 4.1, 2.5],
    'Biomarker3': [3.1, 3.2, 3.3, None, 4.1, 5.1, 3.5]
}

# Create DataFrame
df = pd.DataFrame(data)
df['Date'] = pd.to_datetime(df['Date'])

# Filter data for target patient and location
filtered_df = df[(df['PatientID'] == 244651) & (df['LocationType'] == 'IP')]
filtered_df.set_index('Date', inplace=True)

plt.figure(figsize=(12, 8))

# 关键修改:对每个Biomarker列先删除缺失值,再绘图
for column in filtered_df.columns:
    if column not in ['PatientID', 'LocationType']:
        # 筛选出当前Biomarker非空的行
        valid_data = filtered_df[column].dropna()
        # 使用有效数据的日期索引和值绘图
        plt.plot(valid_data.index, valid_data, marker='o', linestyle='-', label=column)

plt.title('Biomarkers by Date for Patient ID 244651 (IP Location Type)')
plt.xlabel('Date')
plt.ylabel('Biomarker Value')
plt.legend()
plt.grid(True)
plt.xticks(rotation=45)
plt.show()

关键说明

  • filtered_df[column].dropna()会移除当前Biomarker列中值为None的行,只保留有有效数据的日期和对应值。
  • 绘图时直接使用这些有效数据的索引(日期)和值,matplotlib会按日期顺序连接所有有效点,不会插入任何新数据,完全符合需求。

内容的提问来源于stack exchange,提问作者asi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 13:39:56