Python正则findall无法匹配Unicode文本中所有期刊名称问题
Hey there! I see the problem you're hitting—your regex works perfectly on regex101 but only grabs the first match in Python. Let's break this down and fix it.
The Root Cause
Your regex uses the ^ anchor, which by default only matches the start of the entire string in Python's re module. But regex101 automatically enables the Multiline mode for you behind the scenes, which makes ^ match the start of every line instead. That's why it works there but not locally.
The Solution
Add the re.MULTILINE (or re.M for short) flag to your re.findall() call. This tells Python to treat each line's start as a valid position for the ^ anchor.
Here's your updated code:
import re regex = r"^(\d+\)\s*\d+\.\s+)(.*?) ISSN" # Combine re.IGNORECASE with re.MULTILINE using bitwise OR matches = re.findall(regex, txt, re.IGNORECASE | re.MULTILINE) print(len(matches)) print(matches)
Optional Optimization
If you don't need to capture the leading number markers (like 6) 6. ), you can convert that first group into a non-capturing group with (?:...) to clean up your results:
regex = r"^(?:\d+\)\s*\d+\.\s+)(.*?) ISSN" # Now matches will only contain the journal name + frequency matches = re.findall(regex, txt, re.IGNORECASE | re.MULTILINE)
This should now capture all your journal entries exactly like it did on regex101.
内容的提问来源于stack exchange,提问作者chikitin

