如何从URL与文件名拼接的字符串中提取附件的文件名与URL?
解决方案
问题出在URL与下一个文件名无分隔拼接,可利用URL中name=参数值与对应文件名完全一致的特点,编写精准正则来匹配每个文件项:
import re description = "----------------------------------------------\n\nTest Customer, May 11, 2023, 18:27\n\nDo you want to hear a construction joke? Oh, never mind, I'm still working on it!\n\nAttachment(s):\nistockphoto-1131743616-612x612.jpg - https://aaa.zendesk.com/attachments/token/jIUzlZ4ylKke4iP0kB/?name=istockphoto-1131743616-612x612.jpgzendesk_tickets_20230511_091412058891.json - https://aaa.zendesk.com/attachments/token/jIUzlZ4ylKke4iP0kB/?name=zendesk_tickets_20230511_091412058891.json" # 提取附件区域 attachment_section = re.search(r'Attachment\(s\):\n(.+)', description, re.DOTALL).group(1) # 利用URL的name参数与文件名一致的规则,匹配每个文件项 attachments = re.findall(r'([a-zA-Z0-9_-]+\.(jpg|json))\s-\s(https://.*?name=\1)', attachment_section) # 整理结果 attachment_urls = [] for match in attachments: file_name = match[0] url = match[2] attachment_urls.append({ 'file_name': file_name, 'url': url }) # 输出结果 for attachment in attachment_urls: print(f"File Name: {attachment['file_name']}") print(f"URL: {attachment['url']}") print()
代码说明
- 正则
([a-zA-Z0-9_-]+\.(jpg|json))\s-\s(https://.*?name=\1)的核心逻辑:- 捕获第一个分组:匹配以
.jpg或.json结尾的文件名 - 匹配
-分隔符 - 捕获第三个分组:匹配URL,利用
\1反向引用确保URL中name=的参数值与前面捕获的文件名完全一致,自动截断拼接的多余内容
- 捕获第一个分组:匹配以
- 该正则支持任意数量的附件,只要每个附件的URL都符合
name=对应文件名的格式即可。
运行结果
File Name: istockphoto-1131743616-612x612.jpg URL: https://aaa.zendesk.com/attachments/token/jIUzlZ4ylKke4iP0kB/?name=istockphoto-1131743616-612x612.jpg File Name: zendesk_tickets_20230511_091412058891.json URL: https://aaa.zendesk.com/attachments/token/jIUzlZ4ylKke4iP0kB/?name=zendesk_tickets_20230511_091412058891.json
内容的提问来源于stack exchange,提问作者スーパーマン
相关产品推荐
相关产品推荐

