You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python链接爬取正则匹配单一扩展名正常,多扩展名匹配失效

问题解决:正则匹配多格式音频链接仅返回扩展名的问题

问题核心是正则表达式里的捕获组导致re.findall()行为变化:

  • 原正则r'http[s]?://[^\s]+\.mp3'无捕获组,findall返回整个匹配的完整链接;
  • 修改后的r'http[s]?://[^\s]+\.(mp3|flac|wav)'包含捕获组(mp3|flac|wav),findall会优先返回捕获组匹配的内容(即扩展名),而非整个链接。

修正方案

将捕获组改为非捕获组(语法(?:...)),确保findall返回完整匹配链接:

import re
import requests

url = input('Enter URL: ')
html_content = requests.get(url).text

# 使用非捕获组,让findall返回完整链接
pattern = re.compile(r'http[s]?://[^\s]+\.(?:mp3|flac|wav)')

# 提取所有匹配的完整链接
links = re.findall(pattern, html_content)

# 去重
links = list(set(links))

# 写入文件
with open('links.txt', 'w') as file:
    for link in links:
        file.write(link + '\n')

补充说明

若需同时获取完整链接和扩展名,可改用re.finditer()遍历匹配对象:

matches = re.finditer(r'http[s]?://[^\s]+\.(mp3|flac|wav)', html_content)
links = [match.group(0) for match in matches]  # group(0)取完整链接,group(1)取扩展名

内容的提问来源于stack exchange,提问作者Sealfan69

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 07:45:29