You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python读取Clustal格式MSA文件并生成带高亮的对齐PDF

基于Python生成带高亮片段的Clustal格式MSA PDF

前置依赖

先安装两个核心库:

  • Biopython:用于解析Clustal格式的多序列比对文件
  • reportlab:用于生成PDF并实现文本高亮

执行安装命令:

pip install biopython reportlab

步骤1:读取Clustal格式MSA文件

使用Biopython的AlignIO模块读取结构化的比对数据:

from Bio import AlignIO

# 替换为你的Clustal文件路径
alignment = AlignIO.read("your_alignment.clustal", "clustal")

步骤2:定义高亮规则

指定需要高亮的序列ID及其位置区间(注意:位置为0-based,即第一个字符位置为0):

# 示例:键为序列ID,值为(起始位置, 结束位置)的列表(左闭右开)
highlight_regions = {
    "SampleSeq_01": [(12, 28), (35, 47)],
    "SampleSeq_03": [(8, 18)]
}

步骤3:生成带高亮的PDF

通过reportlab将结构化序列数据渲染为PDF,同时对指定区域添加高亮:

from reportlab.lib.pagesizes import letter
from reportlab.lib.styles import getSampleStyleSheet
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer
from reportlab.lib.colors import yellow

def render_msa_pdf(alignment, highlight_config, output_path, line_length=60):
    doc = SimpleDocTemplate(output_path, pagesize=letter)
    styles = getSampleStyleSheet()
    base_style = styles["BodyText"]
    content_elements = []

    for seq_record in alignment:
        seq_id = seq_record.id
        full_seq = str(seq_record.seq)
        
        # 添加加粗的序列ID
        content_elements.append(Paragraph(f"<b>{seq_id}</b>", base_style))
        content_elements.append(Spacer(1, 5))

        # 分段处理长序列,避免超出页面宽度
        for start_idx in range(0, len(full_seq), line_length):
            line_segment = full_seq[start_idx:start_idx+line_length]
            highlighted_segment = ""
            current_global_pos = start_idx

            for char_idx in range(len(line_segment)):
                pos = current_global_pos + char_idx
                # 检查当前位置是否在高亮区间内
                highlight_flag = False
                if seq_id in highlight_config:
                    for (h_start, h_end) in highlight_config[seq_id]:
                        if h_start <= pos < h_end:
                            highlight_flag = True
                            break
                # 拼接带高亮的文本片段
                if highlight_flag:
                    highlighted_segment += f"<span bgColor='{yellow}'>{line_segment[char_idx]}</span>"
                else:
                    highlighted_segment += line_segment[char_idx]
            
            content_elements.append(Paragraph(highlighted_segment, base_style))
        
        content_elements.append(Spacer(1, 10))
    
    doc.build(content_elements)

# 调用函数生成PDF
render_msa_pdf(alignment, highlight_regions, "msa_with_highlight.pdf")

关键说明

  • 为何直接导入PDF无法高亮:直接导出的MSA PDF是静态文本或位图,未保留序列的结构化位置信息,无法通过常规命令定位并修改特定片段样式。必须从原始Clustal文件读取结构化数据,重新渲染为可定制格式的PDF。
  • 位置适配:若习惯使用1-based位置计数,需将高亮区间的起始/结束值各减1(例如1-based的13-29对应0-based的12-28)。
  • 自定义调整:可修改line_length调整每行字符数,替换yellow为lightblue等颜色改变高亮样式,或通过base_style调整字体、字号等排版参数。

内容的提问来源于stack exchange,提问作者Denise Lavezzari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 14:18:10