Python中移除LaTeX字符:去除\textbf{}标签的实现问题
Ah, I see the problem with your current code—your regex is using greedy matching, which causes it to capture far more than just the content inside a single \textbf{} block.
Why Your Original Code Fails
The .* quantifier matches as much text as possible. So when you run re.sub(r'\textbf{(.*)}', '\\1', s), it starts at the first \textbf{ and doesn't stop until it hits the very last } in the entire string. This means it captures everything from 1 + 1} \begin{center} \textbf{some text in here:} \end{center} and \textbf{2 as the group, leading to incorrect output with leftover braces and unprocessed \textbf tags.
The Solution
You need to make the match non-greedy or explicitly match only characters that aren't closing braces. Here are two reliable approaches:
1. Non-Greedy Quantifier (.*?)
Add a ? after .* to make it match the shortest possible string between the opening { and closing }:
import re s = '\textbf{1 + 1} \begin{center} \textbf{some text in here:} \end{center} and \textbf{2} etc' result = re.sub(r'\\textbf{(.*?)}', r'\1', s) print(result)
2. Match Non-Closing Braces ([^}]+)
This approach is more efficient if you know there are no nested braces inside your \textbf{} blocks. It matches any character except }:
result = re.sub(r'\\textbf{([^}]+)}', r'\1', s)
Output
Both methods will produce your desired result:
1 + 1 \begin{center} some text in here: \end{center} and 2 etc
Note on Nested Braces
If you ever need to handle nested braces (e.g., \textbf{text {with} nested braces}), neither of these simple regexes will work. For that, you'd need a recursive regex or a parser specifically designed for TeX, but since your example doesn't involve nested braces, the above solutions are perfect.
内容的提问来源于stack exchange,提问作者dimitris_ps

