如何使用Python实现窗口1000步长20的滑动窗口位点统计与绘图
Python实现方案
依赖库
需要提前安装以下依赖:
pip install pandas matplotlib numpy
完整可运行代码
import pandas as pd import numpy as np import matplotlib.pyplot as plt # 配置参数 window_size = 1000 step_size = 20 # 1. 加载/构造示例数据(实际使用时替换为pd.read_csv读取你的文件即可) data = [ [1250, 1, 1, 1, 1], [1750, np.nan, 1, np.nan, 1], [1786, 1, np.nan, 1, 1], [1812, 1, 1, 1, 1], [1855, 1, 1, 1, 1], [1896, 1, np.nan, 1, np.nan], [2635, np.nan, 1, 1, 1], [1689, 1, 1, np.nan, np.nan], [3250, 1, 1, 1, 1], [3655, 1, np.nan, 1, 1], [3589, np.nan, 1, 1, 1] ] df = pd.DataFrame(data, columns=['POSITION', 'A', 'B', 'C', 'D']) # 位点按位置排序,必须步骤 df = df.sort_values('POSITION').reset_index(drop=True) # 2. 滑动窗口统计 min_pos = df['POSITION'].min() max_pos = df['POSITION'].max() # 生成所有窗口起始位置 window_starts = np.arange(min_pos, max_pos, step_size) result = [] for start in window_starts: end = start + window_size # 筛选当前窗口内的位点 window_df = df[(df['POSITION'] >= start) & (df['POSITION'] < end)] # 统计每个样本有效位点数量(非NA即1) count_a = window_df['A'].count() count_b = window_df['B'].count() count_c = window_df['C'].count() count_d = window_df['D'].count() # 存储结果,用窗口中点作为X轴坐标更直观 result.append([start + window_size/2, count_a, count_b, count_c, count_d]) result_df = pd.DataFrame(result, columns=['window_mid', 'A', 'B', 'C', 'D']) # 3. 绘制趋势图 plt.figure(figsize=(10,6)) # 分别绘制四个样本的曲线 plt.plot(result_df['window_mid'], result_df['A'], label='Sample A', linewidth=1.5) plt.plot(result_df['window_mid'], result_df['B'], label='Sample B', linewidth=1.5) plt.plot(result_df['window_mid'], result_df['C'], label='Sample C', linewidth=1.5) plt.plot(result_df['window_mid'], result_df['D'], label='Sample D', linewidth=1.5) plt.xlabel('Genomic Position (window midpoint)') plt.ylabel('Number of Valid Sites') plt.legend() plt.grid(alpha=0.3) plt.tight_layout() plt.show()
代码说明
- 数据预处理阶段默认会对位点按位置排序,避免原始数据位置乱序导致的统计错误,读取本地文件时使用
pd.read_csv('your_file.txt', sep='\s+', na_values='NA')即可自动识别数据中的NA为缺失值 - 滑动窗口统计逻辑严格匹配需求:窗口大小1000、步长20,统计每个窗口内非NA的有效位点数量
- 绘图参数可根据需求自行调整,包括线条颜色、粗细、坐标轴范围、图片分辨率等
- 数据量较大时可以将窗口统计部分替换为向量化运算提升运行效率,常规量级的数据当前代码运行无压力
内容的提问来源于stack exchange,提问作者user8769986
相关产品推荐
相关产品推荐

