Mac命令行批量处理GenBank文件:提取LOCUS与tRNA信息并分列保存
Batch Processing GenBank Files on Mac: Extract LOCUS and tRNA Entries
Got it, let's work through this batch processing issue for your GenBank files on macOS. I’ve handled similar NCBI data parsing tasks before, so here’s a straightforward, reliable solution that’ll scale to your 2000 files.
The Approach
We’ll create a bash script that:
- Loops through every GenBank file in your target directory
- Extracts the unique
LOCUSidentifier from each file - Pulls all distinct
tRNA-*entries (liketRNA-Glu) - Writes results to a tab-separated values (TSV) file—easy to open in Excel, Numbers, or any data tool
The Script
Create a file named process_genbank.sh with the following code:
#!/bin/bash # Define output file (tab-separated for easy spreadsheet use) OUTPUT_FILE="genbank_results.tsv" # Write header row to output echo -e "FILENAME\tLOCUS\tTRNA_ENTRIES" > "$OUTPUT_FILE" # Loop through all GenBank files (update *.gb to match your file extension, e.g., *.gbk) for FILE in *.gb; do # Skip if the item isn't a file (avoids directories) [ -f "$FILE" ] || continue # Extract the unique LOCUS identifier (first LOCUS line, second field) LOCUS=$(grep -m 1 "^LOCUS" "$FILE" | awk '{print $2}') # Extract all tRNA-* entries, remove duplicates, and format as comma-separated list TRNAS=$(grep -o 'tRNA-[A-Za-z]*' "$FILE" | sort | uniq | tr '\n' ',' | sed 's/,$//') # Handle cases where no tRNA entries exist if [ -z "$TRNAS" ]; then TRNAS="None" fi # Append results to output file echo -e "$FILE\t$LOCUS\t$TRNAS" >> "$OUTPUT_FILE" done echo "Batch processing complete! Results saved to $OUTPUT_FILE"
How to Use
- Move all your GenBank files into a single directory.
- Open Terminal, navigate to that directory:
cd /path/to/your/genbank/files - Create the script:
nano process_genbank.sh(paste the code above, then pressCtrl+Oto save,Ctrl+Xto exit) - Make the script executable:
chmod +x process_genbank.sh - Run the script:
./process_genbank.sh
Key Details & Troubleshooting
- File Extensions: If your GenBank files use
.gbkinstead of.gb, update the*.gbin the script to*.gbk. - LOCUS Extraction: The
grep -m 1 "^LOCUS"ensures we only grab the first (and only) LOCUS line in each file, which is standard for GenBank format.awk '{print $2}'pulls the unique LOCUS ID (e.g.,NZ_AP013294). - tRNA Extraction:
grep -o 'tRNA-[A-Za-z]*'isolates only thetRNA-*strings, avoiding extra text from the GenBank file.sort | uniqremoves any duplicate entries, andtr/sedformats them into a clean comma-separated list. - Performance: 2000 files are no problem for this script—even large GenBank files will process quickly on macOS.
内容的提问来源于stack exchange,提问作者user2861089
相关产品推荐
相关产品推荐

