使用grep和sed优化Bash数字提取脚本的建议咨询
Hey there! Let's dive into optimizing your Bash script for extracting numbers from a target statement and writing them to a file. First, I’ll start with a guess at a common "quick and dirty" temporary solution folks often use (since you didn’t share your exact code):
# Example temporary solution echo "Sample statement with numbers 123, 45.6, and -78" | grep -o '[0-9]*\.[0-9]*\|[0-9]*' > numbers.txt
Here are actionable, practical optimization suggestions tailored to this use case:
Cover all number types robustly
Most basic regex patterns miss edge cases like negative numbers, scientific notation (e.g.,1.2e3), or numbers with leading/trailing whitespace. Expand your regex with extended syntax (using-Eflag for cleaner grouping) to handle these:grep -oE '-?[0-9]+(\.[0-9]+)?([eE][+-]?[0-9]+)?' input.txt > output.txtThis matches negatives, decimals, and scientific notation—way more comprehensive than basic digit matching.
Make the script reusable and configurable
Stop hardcoding your target statement or output file! Turn them into command-line arguments, and add error checking to avoid silent failures:#!/bin/bash # Usage: ./extract_numbers.sh "Your target statement" output_file.txt if [ $# -ne 2 ]; then echo "Error: Missing arguments!" echo "Usage: $0 <target_statement> <output_file>" exit 1 fi echo "$1" | grep -oE '-?[0-9]+(\.[0-9]+)?([eE][+-]?[0-9]+)?' > "$2"Now you can run the script with any input string or file without editing the code every time.
Skip unnecessary
echocalls for file inputs
If your target content lives in a file instead of a string, feed the file directly togrep—it’s faster and avoids extra process overhead:grep -oE '-?[0-9]+(\.[0-9]+)?([eE][+-]?[0-9]+)?' input.txt > output.txtDeduplicate numbers (if needed)
If you want only unique numbers in your output, pipe the results throughsort -uto remove duplicates:grep -oE '-?[0-9]+(\.[0-9]+)?([eE][+-]?[0-9]+)?' input.txt | sort -u > output.txtAdd debugging/verbose mode
For troubleshooting, add a-vflag to print extra context about what the script is doing:#!/bin/bash VERBOSE=false while getopts "v" opt; do case $opt in v) VERBOSE=true ;; *) echo "Invalid option: -$OPTARG" >&2; exit 1 ;; esac done shift $((OPTIND-1)) if [ $# -ne 2 ]; then echo "Usage: $0 [-v] <target_statement> <output_file>" exit 1 fi if $VERBOSE; then echo "Extracting numbers from input: '$1'" echo "Writing results to: $2" fi echo "$1" | grep -oE '-?[0-9]+(\.[0-9]+)?([eE][+-]?[0-9]+)?' > "$2" if $VERBOSE; then echo "Done! Extracted $(wc -l < "$2") numbers total." fiUse
sedfor extra control (alternative approach)
If you need to clean up surrounding non-numeric characters while extracting,sedoffers more flexibility. For example, to split numbers from messy text and filter valid entries:sed -E 's/[^0-9.eE+-]+/ /g; s/^ +| +$//g; s/ +/\n/g' input.txt | grep -E '^-?[0-9]+(\.[0-9]+)?([eE][+-]?[0-9]+)?$' > output.txtThis replaces non-numeric characters with spaces, trims edges, splits into lines, then filters only valid numbers.
内容的提问来源于stack exchange,提问作者moumeneph

