You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python正则findall无法匹配Unicode文本中所有期刊名称问题

Fixing Regex Match Issue in Python

Hey there! I see the problem you're hitting—your regex works perfectly on regex101 but only grabs the first match in Python. Let's break this down and fix it.

The Root Cause

Your regex uses the ^ anchor, which by default only matches the start of the entire string in Python's re module. But regex101 automatically enables the Multiline mode for you behind the scenes, which makes ^ match the start of every line instead. That's why it works there but not locally.

The Solution

Add the re.MULTILINE (or re.M for short) flag to your re.findall() call. This tells Python to treat each line's start as a valid position for the ^ anchor.

Here's your updated code:

import re

regex = r"^(\d+\)\s*\d+\.\s+)(.*?) ISSN"
# Combine re.IGNORECASE with re.MULTILINE using bitwise OR
matches = re.findall(regex, txt, re.IGNORECASE | re.MULTILINE)

print(len(matches))
print(matches)

Optional Optimization

If you don't need to capture the leading number markers (like 6) 6. ), you can convert that first group into a non-capturing group with (?:...) to clean up your results:

regex = r"^(?:\d+\)\s*\d+\.\s+)(.*?) ISSN"
# Now matches will only contain the journal name + frequency
matches = re.findall(regex, txt, re.IGNORECASE | re.MULTILINE)

This should now capture all your journal entries exactly like it did on regex101.

内容的提问来源于stack exchange,提问作者chikitin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:04:36