You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式匹配提取所有格式的13位ISBN编号?

提取多种格式的13位ISBN编号

需要从页面中提取13位ISBN编号,但该编号存在多种不同格式,无法提取所有实例。示例格式如下:

ISBN-13: 978 1 4310 0862 9
ISBN: 9781431008629
ISBN9781431008629
ISBN 9-78-1431-008-629
ISBN: 9781431008629 more text of the number
isbn : 9781431008629 

期望输出格式为:ISBN: 9781431008629

当前使用的Python代码:

myISBN = re.findall("ISBN" + r'[\w\W]{1,17}',text)
myISBN = myISBN[0]
print (myISBN)

改进后的解决方案

通过精准正则匹配前缀、提取纯数字并验证长度,实现统一格式输出:

import re

def extract_isbn(text):
    # 匹配所有ISBN相关片段,不区分大小写
    matches = re.findall(r'(?i)(ISBN(?:-13)?[:\s]?)([\d\s-]+)', text)
    for _, isbn_part in matches:
        # 过滤非数字字符,保留纯数字
        pure_isbn = ''.join(c for c in isbn_part if c.isdigit())
        # 验证是否为13位有效ISBN
        if len(pure_isbn) == 13:
            return f"ISBN: {pure_isbn}"
    return None

# 测试示例
sample_text = """
ISBN-13: 978 1 4310 0862 9
ISBN: 9781431008629
ISBN9781431008629
ISBN 9-78-1431-008-629
ISBN: 9781431008629 more text of the number
isbn : 9781431008629 
"""

print(extract_isbn(sample_text))

关键说明

  • (?i) 让正则不区分大小写,匹配ISBN/isbn等不同写法
  • ISBN(?:-13)?[:\s]? 匹配前缀,支持ISBN、ISBN-13,后接冒号、空格或无分隔符
  • 提取数字部分后过滤掉空格、连字符,只保留纯数字,再验证长度为13位,确保有效性
  • 最终统一输出为ISBN: 纯数字ISBN格式

内容的提问来源于stack exchange,提问作者user1835437

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 14:01:13