You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Mac命令行批量处理GenBank文件:提取LOCUS与tRNA信息并分列保存

Batch Processing GenBank Files on Mac: Extract LOCUS and tRNA Entries

Got it, let's work through this batch processing issue for your GenBank files on macOS. I’ve handled similar NCBI data parsing tasks before, so here’s a straightforward, reliable solution that’ll scale to your 2000 files.

The Approach

We’ll create a bash script that:

  1. Loops through every GenBank file in your target directory
  2. Extracts the unique LOCUS identifier from each file
  3. Pulls all distinct tRNA-* entries (like tRNA-Glu)
  4. Writes results to a tab-separated values (TSV) file—easy to open in Excel, Numbers, or any data tool

The Script

Create a file named process_genbank.sh with the following code:

#!/bin/bash

# Define output file (tab-separated for easy spreadsheet use)
OUTPUT_FILE="genbank_results.tsv"

# Write header row to output
echo -e "FILENAME\tLOCUS\tTRNA_ENTRIES" > "$OUTPUT_FILE"

# Loop through all GenBank files (update *.gb to match your file extension, e.g., *.gbk)
for FILE in *.gb; do
    # Skip if the item isn't a file (avoids directories)
    [ -f "$FILE" ] || continue

    # Extract the unique LOCUS identifier (first LOCUS line, second field)
    LOCUS=$(grep -m 1 "^LOCUS" "$FILE" | awk '{print $2}')

    # Extract all tRNA-* entries, remove duplicates, and format as comma-separated list
    TRNAS=$(grep -o 'tRNA-[A-Za-z]*' "$FILE" | sort | uniq | tr '\n' ',' | sed 's/,$//')
    
    # Handle cases where no tRNA entries exist
    if [ -z "$TRNAS" ]; then
        TRNAS="None"
    fi

    # Append results to output file
    echo -e "$FILE\t$LOCUS\t$TRNAS" >> "$OUTPUT_FILE"
done

echo "Batch processing complete! Results saved to $OUTPUT_FILE"

How to Use

  1. Move all your GenBank files into a single directory.
  2. Open Terminal, navigate to that directory: cd /path/to/your/genbank/files
  3. Create the script: nano process_genbank.sh (paste the code above, then press Ctrl+O to save, Ctrl+X to exit)
  4. Make the script executable: chmod +x process_genbank.sh
  5. Run the script: ./process_genbank.sh

Key Details & Troubleshooting

  • File Extensions: If your GenBank files use .gbk instead of .gb, update the *.gb in the script to *.gbk.
  • LOCUS Extraction: The grep -m 1 "^LOCUS" ensures we only grab the first (and only) LOCUS line in each file, which is standard for GenBank format. awk '{print $2}' pulls the unique LOCUS ID (e.g., NZ_AP013294).
  • tRNA Extraction: grep -o 'tRNA-[A-Za-z]*' isolates only the tRNA-* strings, avoiding extra text from the GenBank file. sort | uniq removes any duplicate entries, and tr/sed formats them into a clean comma-separated list.
  • Performance: 2000 files are no problem for this script—even large GenBank files will process quickly on macOS.

内容的提问来源于stack exchange,提问作者user2861089

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:36:50