You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Julia中使用CSV.read读取存在分隔符缺失的AlphaFold生成PDB文件至DataFrame的解决方案咨询

这个问题其实很常见——AlphaFold输出的PDB偶尔会有这类格式小问题,而且你一开始用空格分隔的思路其实没抓住PDB文件的核心本质:PDB的ATOM行是固定列宽格式,不是靠空格分隔的!下面给你几个比逐行手动补空格更优雅、更可靠的解决方案:

解决方案1:用专业生物信息学包(最推荐)

BioJulia生态里的BioStructures.jl是专门处理PDB、mmCIF这类生物结构文件的工具,它会自动处理各种PDB格式的边缘情况(包括你遇到的分隔符缺失问题),还能直接提取原子的生物属性。步骤如下:

  1. 先安装所需包:
using Pkg
Pkg.add("BioStructures")
Pkg.add("DataFrames")
  1. 读取PDB并转换为DataFrame:
using BioStructures, DataFrames

# 读取PDB文件,过滤出标准ATOM行
struc = read("your_file.pdb", PDB)
atoms = collectatoms(struc, standardselector)

# 映射为你需要的DataFrame格式
df = DataFrame(
    Type = fill("ATOM", length(atoms)),
    Index = [a.number for a in atoms],
    Atom = [a.name for a in atoms],
    Amino_acid_type = [resname(a) for a in atoms],
    Chain = [chainid(a) for a in atoms],
    Amino_acid_number = [resnumber(a) for a in atoms],
    Position_X = [coords(a)[1] for a in atoms],
    Position_Y = [coords(a)[2] for a in atoms],
    Position_Z = [coords(a)[3] for a in atoms],
    Something = [a.occupancy for a in atoms],
    pIDDT = [a.bfactor for a in atoms],  # AlphaFold的pIDDT值存储在bfactor字段中
    Atom_type = [element(a) for a in atoms]
)

这个方法完全不用自己处理格式异常,专业包已经覆盖了所有PDB规范和常见的输出问题,是最省心的选择。

解决方案2:按固定列宽手动解析(无需专业包)

如果你不想引入生物信息学依赖,可以直接利用PDB的固定列宽规则读取。根据PDB官方规范,ATOM行的关键字段位置是固定的:

  • Type: 1-6
  • Index: 7-11
  • Atom: 13-16
  • Amino acid type: 18-20
  • Chain: 22-22
  • Amino acid number: 23-26
  • Position X: 31-38
  • Position Y: 39-46
  • Position Z: 47-54
  • Occupancy(你的Something): 55-60
  • B-factor(你的pIDDT): 61-66
  • Atom type: 77-78

用CSV.read的columnspecs参数指定每个字段的起止位置即可:

using CSV, DataFrames

headers = ["Type", "Index", "Atom", "Amino acid type", "Chain", "Amino acid number", "Position X", "Position Y", "Position Z", "Something", "pIDDT", "Atom type"]
# 每个字段的(start, end)位置(Julia是1-based索引)
colspecs = [
    (1,6), (7,11), (13,16), (18,20), (22,22), (23,26),
    (31,38), (39,46), (47,54), (55,60), (61,66), (77,78)
]
types = Dict(
    "Type"=>String, "Index"=>Int64, "Atom"=>String, "Amino acid type"=>String,
    "Chain"=>Char, "Amino acid number"=>Int64,
    "Position X"=>Float64, "Position Y"=>Float64, "Position Z"=>Float64,
    "Something"=>Float64, "pIDDT"=>Float64, "Atom type"=>Char
)

df = CSV.read(
    "your_file.pdb", DataFrame;
    header=headers,
    skipto=2,  # 跳过开头的MODEL行
    delim=nothing,
    columnspecs=colspecs,
    types=types,
    trimwhitespace=true  # 自动去除字段前后的空格
)

这种方法直接按规范截取字段,不管行内有没有缺失的空格,都能准确提取数据,比空格分隔的方式可靠得多。

解决方案3:高效批量预处理(坚持用空格分隔时)

如果你还是想沿用空格分隔的思路,可以用正则表达式批量修复所有缺失空格的负号,比逐行判断高效:

using CSV, DataFrames

# 读取所有行,用正则修复格式
lines = readlines("your_file.pdb")
# 匹配数字后紧跟的负号,替换成" 负号"
fixed_lines = replace.(lines, r"(?<=\d)-" => " -")

# 用IOBuffer避免写入临时文件
io = IOBuffer(join(fixed_lines, "\n"))
headers = ["Type", "Index", "Atom", "Amino acid type", "Chain", "Amino acid number", "Position X", "Position Y", "Position Z", "Something", "pIDDT", "Atom type"]
types = Dict("Type"=>String, "Index"=>Int64, "Atom"=>String, "Amino acid type"=>String, "Chain"=>Char, "Amino acid number"=>Int64, "Position X"=>Float64, "Position Y"=>Float64, "Position Z"=>Float64, "Something"=>Float64, "pIDDT"=>Float64, "Atom type"=>Char)

df = CSV.read(io, DataFrame; header=headers, skipto=2, delim=' ', ignorerepeated=true, types=types)

这里的正则(?<=\d)-会精准定位所有前面是数字的负号(比如24.512-100.624里的-),批量补上缺失的空格,再用你原来的逻辑读取即可。

内容的提问来源于stack exchange,提问作者Aephir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 00:02:50