Julia中使用CSV.read读取存在分隔符缺失的AlphaFold生成PDB文件至DataFrame的解决方案咨询
这个问题其实很常见——AlphaFold输出的PDB偶尔会有这类格式小问题,而且你一开始用空格分隔的思路其实没抓住PDB文件的核心本质:PDB的ATOM行是固定列宽格式,不是靠空格分隔的!下面给你几个比逐行手动补空格更优雅、更可靠的解决方案:
解决方案1:用专业生物信息学包(最推荐)
BioJulia生态里的BioStructures.jl是专门处理PDB、mmCIF这类生物结构文件的工具,它会自动处理各种PDB格式的边缘情况(包括你遇到的分隔符缺失问题),还能直接提取原子的生物属性。步骤如下:
- 先安装所需包:
using Pkg Pkg.add("BioStructures") Pkg.add("DataFrames")
- 读取PDB并转换为DataFrame:
using BioStructures, DataFrames # 读取PDB文件,过滤出标准ATOM行 struc = read("your_file.pdb", PDB) atoms = collectatoms(struc, standardselector) # 映射为你需要的DataFrame格式 df = DataFrame( Type = fill("ATOM", length(atoms)), Index = [a.number for a in atoms], Atom = [a.name for a in atoms], Amino_acid_type = [resname(a) for a in atoms], Chain = [chainid(a) for a in atoms], Amino_acid_number = [resnumber(a) for a in atoms], Position_X = [coords(a)[1] for a in atoms], Position_Y = [coords(a)[2] for a in atoms], Position_Z = [coords(a)[3] for a in atoms], Something = [a.occupancy for a in atoms], pIDDT = [a.bfactor for a in atoms], # AlphaFold的pIDDT值存储在bfactor字段中 Atom_type = [element(a) for a in atoms] )
这个方法完全不用自己处理格式异常,专业包已经覆盖了所有PDB规范和常见的输出问题,是最省心的选择。
解决方案2:按固定列宽手动解析(无需专业包)
如果你不想引入生物信息学依赖,可以直接利用PDB的固定列宽规则读取。根据PDB官方规范,ATOM行的关键字段位置是固定的:
- Type: 1-6
- Index: 7-11
- Atom: 13-16
- Amino acid type: 18-20
- Chain: 22-22
- Amino acid number: 23-26
- Position X: 31-38
- Position Y: 39-46
- Position Z: 47-54
- Occupancy(你的Something): 55-60
- B-factor(你的pIDDT): 61-66
- Atom type: 77-78
用CSV.read的columnspecs参数指定每个字段的起止位置即可:
using CSV, DataFrames headers = ["Type", "Index", "Atom", "Amino acid type", "Chain", "Amino acid number", "Position X", "Position Y", "Position Z", "Something", "pIDDT", "Atom type"] # 每个字段的(start, end)位置(Julia是1-based索引) colspecs = [ (1,6), (7,11), (13,16), (18,20), (22,22), (23,26), (31,38), (39,46), (47,54), (55,60), (61,66), (77,78) ] types = Dict( "Type"=>String, "Index"=>Int64, "Atom"=>String, "Amino acid type"=>String, "Chain"=>Char, "Amino acid number"=>Int64, "Position X"=>Float64, "Position Y"=>Float64, "Position Z"=>Float64, "Something"=>Float64, "pIDDT"=>Float64, "Atom type"=>Char ) df = CSV.read( "your_file.pdb", DataFrame; header=headers, skipto=2, # 跳过开头的MODEL行 delim=nothing, columnspecs=colspecs, types=types, trimwhitespace=true # 自动去除字段前后的空格 )
这种方法直接按规范截取字段,不管行内有没有缺失的空格,都能准确提取数据,比空格分隔的方式可靠得多。
解决方案3:高效批量预处理(坚持用空格分隔时)
如果你还是想沿用空格分隔的思路,可以用正则表达式批量修复所有缺失空格的负号,比逐行判断高效:
using CSV, DataFrames # 读取所有行,用正则修复格式 lines = readlines("your_file.pdb") # 匹配数字后紧跟的负号,替换成" 负号" fixed_lines = replace.(lines, r"(?<=\d)-" => " -") # 用IOBuffer避免写入临时文件 io = IOBuffer(join(fixed_lines, "\n")) headers = ["Type", "Index", "Atom", "Amino acid type", "Chain", "Amino acid number", "Position X", "Position Y", "Position Z", "Something", "pIDDT", "Atom type"] types = Dict("Type"=>String, "Index"=>Int64, "Atom"=>String, "Amino acid type"=>String, "Chain"=>Char, "Amino acid number"=>Int64, "Position X"=>Float64, "Position Y"=>Float64, "Position Z"=>Float64, "Something"=>Float64, "pIDDT"=>Float64, "Atom type"=>Char) df = CSV.read(io, DataFrame; header=headers, skipto=2, delim=' ', ignorerepeated=true, types=types)
这里的正则(?<=\d)-会精准定位所有前面是数字的负号(比如24.512-100.624里的-),批量补上缺失的空格,再用你原来的逻辑读取即可。
内容的提问来源于stack exchange,提问作者Aephir
相关产品推荐
相关产品推荐

