You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何截取PLINK fam文件第一列ID前两段,保留其余列内容?

需求

PLINK的FAM文件中,样本ID(第二列)与家系ID(第一列)格式均为Continent_Breed_Ind-ID,需要将第一列修改为仅保留Continent_Breed部分,同时完整保留其余列内容。

模拟数据集

Continent1_Breed1_Ind-ID1 Continent1_Breed1_Ind-ID1 0 0 0 -9
Continent1_Breed2_Ind-ID2 Continent1_Breed2_Ind-ID1 0 0 0 -0
Continent2_Breed1_Ind-ID1 Continent2_Breed1_Ind-ID1 0 0 0 -9

期望结果

Continent1_Breed1 Continent1_Breed1_Ind-ID1 0 0 0 -9
Continent1_Breed2 Continent1_Breed2_Ind-ID1 0 0 0 -0
Continent2_Breed1 Continent2_Breed1_Ind-ID1 0 0 0 -9

错误尝试

使用以下sed命令时,仅返回第一列,无法保留其余内容:

sed -r 's/_[^_]*//2g' file.fam

解决方案

方法1:使用sed精准修改第一列

通过正则仅匹配第一列中第二个下划线之后的内容并替换,不影响后续列:

sed -r 's/^([^_]+_[^_]+)_[^ ]+/\1/' file.fam
  • 正则解释:
    • ^([^_]+_[^_]+):捕获第一列中前两个下划线分隔的部分(即Continent_Breed)
    • _[^ ]+:匹配第一列中第二个下划线到第一个空格之间的所有内容(即_Ind-IDxxx)
    • \1:将匹配到的内容替换为之前捕获的Continent_Breed部分

方法2:使用awk更直观处理

通过分割第一列提取目标内容,重新赋值后输出整行:

awk '{split($1, arr, "_"); $1 = arr[1] "_" arr[2]; print}' file.fam
  • 逻辑解释:
    • split($1, arr, "_"):将第一列按下划线分割为数组arr
    • $1 = arr[1] "_" arr[2]:将第一列重新赋值为数组前两个元素的拼接结果
    • print:输出修改后的整行内容

内容的提问来源于stack exchange,提问作者Pedro Morell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 12:45:37