Bash批量处理文件名:提取核心标识及配对文件匹配问题
Hey there! Let's break down why your sed command isn't giving you the expected core identifiers from your MP3 filenames.
The Root Cause
When you tested your regex on regex101, you were likely using Extended Regular Expressions (ERE)—but by default, sed uses Basic Regular Expressions (BRE). The problem is that meta-characters like + (one or more matches) and | (alternation) are treated as literal characters in BRE, not as regex operators. That's why your substitution wasn't working, and you ended up with the original filenames instead of the trimmed core identifiers.
Fixing the Command
You have two straightforward ways to get this working:
1. Use Extended Regular Expressions with Sed
For GNU sed (most Linux systems), add the -r flag to enable ERE. For BSD sed (common on macOS), use -E instead. Also, note that you should escape the leading . in your regex (to match the literal dot before the file extension) and add a $ to anchor the match to the end of the filename:
# GNU sed ls -1 *.mp3 | sed -r 's/\.[^_\d|\W]+\d?$//' > list_mp3 # BSD sed (macOS) ls -1 *.mp3 | sed -E 's/\.[^_\d|\W]+\d?$//' > list_mp3
2. Escape Meta-Characters for Basic Regular Expressions
If you can't use ERE flags, you can escape the + and ? in your regex to give them special meaning in BRE:
ls -1 *.mp3 | sed 's/\.[^_\d|\W]\+\d\?//' > list_mp3
This works, but it's less readable than using ERE, so the first option is preferred when possible.
Quick Verification
Let's say you have a filename like recording_25-09-2018 00-08-33 (1).mp3. The corrected command should output the core identifier you're expecting (like _25-09-2018 00-08-33 (1) or the full prefix+timestamp, depending on your actual filename structure). This will let you easily cross-reference with your .txt files to find paired records.
内容的提问来源于stack exchange,提问作者Eugene Mikhalov

