如何在存在缺失值时用线条连接Biomarker数据点
问题描述
需要在同一张图中绘制多个Biomarker随日期的变化曲线,但各Biomarker的采样日期不同且存在缺失值。当前绘图代码因缺失值导致线条断裂,要求直接连接同一Biomarker的所有有效数据点,且禁止使用interpolate生成新数据点。
示例数据:
data = { 'PatientID': [244651, 244651, 244651, 244651, 244652, 244653, 244651], 'LocationType': ['IP', 'IP', 'OP', 'IP', 'IP', 'OP', 'IP'], 'Date': ['2023-01-01', '2023-01-02', '2023-01-03', '2023-01-04', '2023-01-01', '2023-01-01', '2023-01-05'], 'Biomarker1': [1.1, 1.2, None, 1.4, 2.1, 3.1, 1.5], 'Biomarker2': [2.1, None, 2.3, 2.4, 3.1, 4.1, 2.5], 'Biomarker3': [3.1, 3.2, 3.3, None, 4.1, 5.1, 3.5] }
原完整代码:
import pandas as pd import matplotlib.pyplot as plt import numpy as np # Sample DataFrame (replace this with your actual DataFrame) data = { 'PatientID': [244651, 244651, 244651, 244651, 244652, 244653, 244651], 'LocationType': ['IP', 'IP', 'OP', 'IP', 'IP', 'OP', 'IP'], 'Date': ['2023-01-01', '2023-01-02', '2023-01-03', '2023-01-04', '2023-01-01', '2023-01-01', '2023-01-05'], 'Biomarker1': [1.1, 1.2, None, 1.4, 2.1, 3.1, 1.5], 'Biomarker2': [2.1, None, 2.3, 2.4, 3.1, 4.1, 2.5], 'Biomarker3': [3.1, 3.2, 3.3, None, 4.1, 5.1, 3.5] } # Create DataFrame df = pd.DataFrame(data) df['Date'] = pd.to_datetime(df['Date']) # Filter the data for the specified patient ID and IP location type filtered_df = df[(df['PatientID'] == 244651) & (df['LocationType'] == 'IP')] # Set the date as the index filtered_df.set_index('Date', inplace=True) # Plot all biomarkers plt.figure(figsize=(12, 8)) # Loop through each biomarker column to plot each one separately for column in filtered_df.columns: if column not in ['PatientID', 'LocationType']: plt.plot(filtered_df.index, filtered_df[column], marker='o', linestyle='-', label=column) plt.title('Biomarkers by Date for Patient ID 244651 (IP Location Type)') plt.xlabel('Date') plt.ylabel('Biomarker Value') plt.legend() plt.grid(True) plt.xticks(rotation=45) plt.show()
解决方案
核心方法是对每个Biomarker列单独筛选出无缺失值的行,仅保留有效数据点后再绘图,这样matplotlib会直接连接这些有效点,不会因缺失值断裂,也不会生成新数据。
修改后的完整代码:
import pandas as pd import matplotlib.pyplot as plt import numpy as np # Sample DataFrame data = { 'PatientID': [244651, 244651, 244651, 244651, 244652, 244653, 244651], 'LocationType': ['IP', 'IP', 'OP', 'IP', 'IP', 'OP', 'IP'], 'Date': ['2023-01-01', '2023-01-02', '2023-01-03', '2023-01-04', '2023-01-01', '2023-01-01', '2023-01-05'], 'Biomarker1': [1.1, 1.2, None, 1.4, 2.1, 3.1, 1.5], 'Biomarker2': [2.1, None, 2.3, 2.4, 3.1, 4.1, 2.5], 'Biomarker3': [3.1, 3.2, 3.3, None, 4.1, 5.1, 3.5] } # Create DataFrame df = pd.DataFrame(data) df['Date'] = pd.to_datetime(df['Date']) # Filter data for target patient and location filtered_df = df[(df['PatientID'] == 244651) & (df['LocationType'] == 'IP')] filtered_df.set_index('Date', inplace=True) plt.figure(figsize=(12, 8)) # 关键修改:对每个Biomarker列先删除缺失值,再绘图 for column in filtered_df.columns: if column not in ['PatientID', 'LocationType']: # 筛选出当前Biomarker非空的行 valid_data = filtered_df[column].dropna() # 使用有效数据的日期索引和值绘图 plt.plot(valid_data.index, valid_data, marker='o', linestyle='-', label=column) plt.title('Biomarkers by Date for Patient ID 244651 (IP Location Type)') plt.xlabel('Date') plt.ylabel('Biomarker Value') plt.legend() plt.grid(True) plt.xticks(rotation=45) plt.show()
关键说明
filtered_df[column].dropna()会移除当前Biomarker列中值为None的行,只保留有有效数据的日期和对应值。- 绘图时直接使用这些有效数据的索引(日期)和值,matplotlib会按日期顺序连接所有有效点,不会插入任何新数据,完全符合需求。
内容的提问来源于stack exchange,提问作者asi
相关产品推荐
相关产品推荐

