如何在Bash中提取文件内指定字符串并关联对应文件名
问题描述
我有数千个分别存放在各自文件夹中的文件,需要将特定字符串ID(以"Sample"开头)与对应的文件名关联起来。示例文件夹结构如下:
Folder_1 ├── file_abc.txt └── Sample-001-abc Folder_2 ├── file_efg.txt └── Sample-002-efg Folder_3 ├── file_hig.txt └── Sample-003-hig
我需要基于"Sample"字符串进行匹配,将其与所在文件名关联,确保顺序对应,最终输出文件Filename_sample_linked.txt,内容格式如下:
file_abc.txt Sample-001-abc file_efg.txt Sample-002-efg file_hig.txt Sample-003-hig ...etc
目前我用以下命令获取所有ID,但不知道如何将ID和来源文件名对应:
find dir/*.txt -type f -exec grep -H 'Sample' {} + >> List_of_samples.txt
解决方案
方法1:直接处理grep的输出
你用的grep -H命令已经会输出文件名:匹配到的Sample字符串格式的内容(比如Folder_1/file_abc.txt:Sample-001-abc),只需要做简单的格式转换就能得到目标结果:
用sed处理
find dir -name "*.txt" -type f -exec grep -H '^Sample' {} + | sed -E 's|.*/([^/:]+):(Sample.*)|\1 \2|' > Filename_sample_linked.txt
^Sample确保只匹配以Sample开头的行(避免文件中其他位置出现的无关Sample字符串)- sed命令的作用:去掉文件名的路径前缀,把文件名和匹配到的ID用空格分隔
用awk处理
find dir -name "*.txt" -type f -exec grep -H '^Sample' {} + | awk -F: '{sub(/.*\//, "", $1); print $1, $2}' > Filename_sample_linked.txt
-F:以冒号为分隔符拆分每行内容sub(/.*\//, "", $1)去掉文件名中的路径部分,只保留纯文件名print $1, $2输出文件名和Sample ID,用空格分隔
方法2:遍历文件直接提取
如果每个txt文件里只有一个需要匹配的Sample ID,也可以直接遍历文件读取对应行:
> Filename_sample_linked.txt # 先清空目标文件 find dir -name "*.txt" -type f | while read -r file; do sample_id=$(grep '^Sample' "$file") if [ -n "$sample_id" ]; then echo "$(basename "$file") $sample_id" >> Filename_sample_linked.txt fi done
basename "$file"获取不带路径的纯文件名- 先判断
sample_id非空再输出,避免生成空行
补充说明
- 如果Sample字符串不一定在行首,去掉命令中的
^即可 - 如果单个文件里有多个Sample ID,可添加
grep -m1参数只取第一个匹配项,避免一行对应多个ID的情况
内容的提问来源于stack exchange,提问作者vanish007
相关产品推荐
相关产品推荐

