如何用VBA正则表达式从漫画标题字符串中提取期号
问题
我不擅长正则表达式,需要从包含漫画名称、有时带作者名的长标题字符串里提取期号。期号一般是字符串里最后一个数字,但也有例外。以下是6种典型变体示例:
STAR WARS: DOCTOR APHRA 32 CHRIS SPROUSE RETURN OF THE JEDI 40TH ANNIVERSARY VARIANT
DEADPOOL 7
X-23: DEADLY REGENESIS 3 GERALD PAREL VARIANT
SPIDER-MAN 2099: DARK GENESIS 5
THE GODFORSAKEN 99 OF KRONOS 2 KEN GRAGGINS VARIANT
Teenage Mutant Ninja Turtles: Saturday Morning Adventures (2023-) #1 Variant RI (10) (Dooney)
我现在用VBA写了一个函数,但它匹配不了Spider-Man 2099这种情况,还可能漏掉其他变体。另外这个函数需要获取匹配位置用于其他操作,希望通过特定模式序列逐步匹配,尽可能精准。现有函数代码如下:
Function ExtractText(c As Range) As String Dim rgx As RegExp Dim match As match Dim mc As MatchCollection Dim sComicNo As String, sPattern As Variant Dim lPos As Long, x As Long sComicNo = "" sPattern = Array(" [0-9] ", " #[0-9] ", " [0-9][0-9] ", " #[0-9][0-9] ", " [0-9][0-9]", " #[0-9][0-9]", " [0-9]", " #[0-9]") lPos = 0 Set rgx = New RegExp On Error GoTo ErrHandler Do While sComicNo = "" With rgx .Pattern = sPattern(x) .Global = True If .Test(c.Value) Then Set mc = .Execute(c.Value) If mc.Count > 0 Then Set match = mc.Item(mc.Count - 1) Else ExtractText = "" End If lPos = match.FirstIndex sComicNo = WorksheetFunction.Trim(match.Value) & "|" & lPos Else sComicNo = "" End If x = x + 1 If x > 8 Then ExtractText = sComicNo Exit Function End If End With Loop ExtractText = sComicNo ErrHandler: Exit Function End Function
解决方案
核心问题是原模式未区分系列编号/年份和期号,比如Spider-Man 2099里的2099是系列编号而非期号。我们可以调整匹配逻辑,按优先级尝试精准模式:
优化思路
- 优先匹配带
#的期号(比如#1、#10),这是最明确的期号标识 - 匹配字符串末尾的纯数字,处理无
#的基础情况 - 排除带年份后缀(如40TH、20TH)的数字,避免误判年份为期刊号
改进后的VBA函数
Function ExtractIssueNumber(c As Range) As String Dim rgx As RegExp Dim match As match Dim mc As MatchCollection Dim sPatterns As Variant Dim x As Long Dim issueNo As String Dim pos As Long ' 按优先级排序的正则模式:越靠前匹配精度越高 sPatterns = Array( _ "#(\d+)(?!.*#\d+)", ' 匹配最后一个带#的数字(捕获数字部分) "\b(\d+)\s*$", ' 匹配字符串末尾的纯数字 "\b(\d+)\s+(?!\d+TH|\d+ST|\d+ND|\d+RD)[A-Za-z]+" ' 匹配后跟非年份后缀的数字 ) Set rgx = New RegExp rgx.Global = True issueNo = "" pos = -1 For x = LBound(sPatterns) To UBound(sPatterns) rgx.Pattern = sPatterns(x) If rgx.Test(c.Value) Then Set mc = rgx.Execute(c.Value) ' 取最后一个匹配项(确保是最靠后的期号) Set match = mc(mc.Count - 1) ' 提取捕获组中的纯数字,去掉#前缀 issueNo = match.SubMatches(0) ' 调整位置:如果是带#的匹配,位置跳过#的占位 pos = match.FirstIndex + IIf(match.Value Like "#*", 1, 0) Exit For End If Next x ' 返回「期号|位置」格式的结果 If issueNo <> "" Then ExtractIssueNumber = issueNo & "|" & pos Else ExtractIssueNumber = "" End If End Function
模式说明
#(\d+)(?!.*#\d+):精准匹配最后一个带#的数字,确保后面没有其他带#的数字,避免误匹配中间的编号\b(\d+)\s*$:匹配字符串末尾的纯数字,处理如DEADPOOL 7这类简单场景\b(\d+)\s+(?!\d+TH|\d+ST|\d+ND|\d+RD)[A-Za-z]+:匹配后跟普通英文单词的数字,排除40TH这类年份后缀的数字
测试结果
针对你提供的6个示例,提取结果如下:
32|257|73|265|302|331|64
内容的提问来源于stack exchange,提问作者Starnes Student
相关产品推荐
相关产品推荐

