OSX与Linux下Grep差异:VCF位点提取脚本跨平台失效排查
Here are actionable improvements to make your script robust, efficient, and cross-platform compatible:
1. Use Fixed-String Matching (Eliminate Regex Ambiguity)
Avoid relying on inconsistent regex behavior by forcing grep to treat your site strings as fixed literals with the -F flag (equivalent to fgrep). This works the same way on both BSD and GNU grep:
grep -F "$line" "$vcf"
This prevents accidental parsing of special characters (like \, ., or *) in your site strings as regex metacharacters.
2. Batch Matching (Massive Performance Boost)
Looping through each line and running grep repeatedly is inefficient, especially for large VCF or sites files. Use grep’s -f flag to load all match patterns from your sites file in one pass—this reads the VCF only once instead of once per line:
# Keep VCF header lines (optional but recommended) grep "^#" "$2" > "$3" # Match all sites in a single pass and append to output grep -F -f "$1" "$2" >> "$3"
This is orders of magnitude faster for large datasets.
3. Enforce Proper Tab Characters in Input
Ensure your sites file uses actual tabs instead of literal \t strings. If you need to convert an existing file with \t to real tabs, use this cross-platform sed command:
sed 's/\\t/\t/g' old_sites.txt > new_sites.txt
4. Add Input Validation
Make your script user-friendly by checking for missing arguments or non-existent files:
#!/bin/bash if [ $# -ne 3 ]; then echo "Usage: $0 <sites_file> <input_vcf> <output_vcf>" >&2 exit 1 fi sites_file="$1" input_vcf="$2" output_vcf="$3" # Verify input files exist for file in "$sites_file" "$input_vcf"; do if [ ! -f "$file" ]; then echo "Error: File '$file' does not exist." >&2 exit 1 fi done # Run extraction grep "^#" "$input_vcf" > "$output_vcf" grep -F -f "$sites_file" "$input_vcf" >> "$output_vcf" echo "Extraction complete. Output saved to $output_vcf"
5. Optional: Handle Chromosome Naming Variants
Some VCFs use chr1 instead of Chr01 or 1—if this is a risk, normalize chromosome names in both your sites file and VCF first. For example, convert all to lowercase chr prefixes:
# Normalize sites file sed 's/^Chr/chr/' "$sites_file" > normalized_sites.txt # Normalize VCF and extract matches grep "^#" "$input_vcf" > "$output_vcf" grep -F -f normalized_sites.txt <(sed 's/^Chr/chr/' "$input_vcf") >> "$output_vcf"
内容的提问来源于stack exchange,提问作者C. John

