如何用sed高效去除美元符号单侧空格?支持跨行$$场景
Let's tackle this step by step—your initial attempts were close, but greedy matching and line-by-line processing were tripping you up. Here's how to fix both inline $...$ formulas and multiline $$...$$ blocks correctly:
1. Fix Inline $...$ Formulas
The problem with your first capture group command was that .* greedily matches everything, including spaces and other $ characters. Instead, use [^$]* to match only characters that aren't a dollar sign, which ensures we stop at the closing $.
Run this GNU sed command:
sed -E 's/\$\s*([^$]*)\s*\$/\$\1\$/g' input.txt
Breakdown:
\$\s*: Matches an opening$followed by any number of spaces/tabs([^$]*): Captures all characters except$(so it stops at the closing delimiter)\s*\$: Matches any number of spaces/tabs before the closing$\$\1\$: Replaces with the opening$, captured content, and closing$(no extra spaces)
Testing your example:
Input:
Some text $ latex code $ some more text
Output:Some text $latex code$ some more text
This won't touch standalone $ characters like $ some because they don't have a matching closing $ in the same line.
2. Fix Multiline $$...$$ Blocks
For block formulas that span multiple lines, sed's default line-by-line processing won't work. Use GNU sed's -z flag to read the entire file as a single stream, and adjust the pattern to target $$ delimiters:
sed -zE 's/\$\$\s*([^\$]*?)\s*\$\$/\$\$\1\$\$/g' input.txt
Breakdown:
-z: Treats the input as a null-terminated stream (so we can match across lines)\$\$\s*: Matches opening$$followed by any whitespace([^\$]*?): Non-greedy capture of all characters except$(stops at the first closing$$)\s*\$\$: Matches whitespace before the closing$$
This works even if your block formula looks like:
$$ \begin{align} E = mc^2 \end{align} $$
It will trim the spaces around the content to:
$$ \begin{align} E = mc^2 \end{align} $$
3. Combine Both Fixes in One Command
To handle both inline and block formulas at once, chain the two substitutions:
sed -zE 's/\$\$\s*([^\$]*?)\s*\$\$/\$\$\1\$\$/g; s/\$\s*([^$]*)\s*\$/\$\1\$/g' input.txt
Notes:
- This uses GNU sed features (
-E,-z). If you're on macOS, install GNU sed via Homebrew (brew install gnu-sed) and usegsedinstead ofsed. - If your LaTeX code contains escaped
\$characters (literal dollar signs), you'll need to adjust the regex to ignore those—let me know if that's a case you need to handle!
内容的提问来源于stack exchange,提问作者physicophilic

