You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python实现窗口1000步长20的滑动窗口位点统计与绘图

Python实现方案

依赖库

需要提前安装以下依赖:

pip install pandas matplotlib numpy

完整可运行代码

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt

# 配置参数
window_size = 1000
step_size = 20

# 1. 加载/构造示例数据(实际使用时替换为pd.read_csv读取你的文件即可)
data = [
    [1250, 1, 1, 1, 1],
    [1750, np.nan, 1, np.nan, 1],
    [1786, 1, np.nan, 1, 1],
    [1812, 1, 1, 1, 1],
    [1855, 1, 1, 1, 1],
    [1896, 1, np.nan, 1, np.nan],
    [2635, np.nan, 1, 1, 1],
    [1689, 1, 1, np.nan, np.nan],
    [3250, 1, 1, 1, 1],
    [3655, 1, np.nan, 1, 1],
    [3589, np.nan, 1, 1, 1]
]
df = pd.DataFrame(data, columns=['POSITION', 'A', 'B', 'C', 'D'])
# 位点按位置排序,必须步骤
df = df.sort_values('POSITION').reset_index(drop=True)

# 2. 滑动窗口统计
min_pos = df['POSITION'].min()
max_pos = df['POSITION'].max()
# 生成所有窗口起始位置
window_starts = np.arange(min_pos, max_pos, step_size)
result = []

for start in window_starts:
    end = start + window_size
    # 筛选当前窗口内的位点
    window_df = df[(df['POSITION'] >= start) & (df['POSITION'] < end)]
    # 统计每个样本有效位点数量(非NA即1)
    count_a = window_df['A'].count()
    count_b = window_df['B'].count()
    count_c = window_df['C'].count()
    count_d = window_df['D'].count()
    # 存储结果,用窗口中点作为X轴坐标更直观
    result.append([start + window_size/2, count_a, count_b, count_c, count_d])

result_df = pd.DataFrame(result, columns=['window_mid', 'A', 'B', 'C', 'D'])

# 3. 绘制趋势图
plt.figure(figsize=(10,6))
# 分别绘制四个样本的曲线
plt.plot(result_df['window_mid'], result_df['A'], label='Sample A', linewidth=1.5)
plt.plot(result_df['window_mid'], result_df['B'], label='Sample B', linewidth=1.5)
plt.plot(result_df['window_mid'], result_df['C'], label='Sample C', linewidth=1.5)
plt.plot(result_df['window_mid'], result_df['D'], label='Sample D', linewidth=1.5)

plt.xlabel('Genomic Position (window midpoint)')
plt.ylabel('Number of Valid Sites')
plt.legend()
plt.grid(alpha=0.3)
plt.tight_layout()
plt.show()

代码说明

  • 数据预处理阶段默认会对位点按位置排序,避免原始数据位置乱序导致的统计错误,读取本地文件时使用pd.read_csv('your_file.txt', sep='\s+', na_values='NA')即可自动识别数据中的NA为缺失值
  • 滑动窗口统计逻辑严格匹配需求:窗口大小1000、步长20,统计每个窗口内非NA的有效位点数量
  • 绘图参数可根据需求自行调整,包括线条颜色、粗细、坐标轴范围、图片分辨率等
  • 数据量较大时可以将窗口统计部分替换为向量化运算提升运行效率,常规量级的数据当前代码运行无压力

内容的提问来源于stack exchange,提问作者user8769986

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 14:24:02