You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中忽略URL前缀后缀,匹配关键词并输出完整URL

解决URL格式不一致的匹配问题

你的两次尝试都卡在了URL格式差异上——要么只能精确匹配完全一致的字符串,要么匹配结果跟着fileURLs的格式变来变去。核心问题是没消除www.、http://、https://这些前缀的干扰,我们需要先把所有URL标准化处理,再进行匹配。

第一步:编写URL标准化函数

这个函数会把各种格式的URL转换成统一形式,去掉无关前缀、统一小写:

def normalize_url(url):
    # 转小写,避免大小写差异
    normalized = str(url).lower()
    # 移除http://或https://前缀
    if normalized.startswith(('http://', 'https://')):
        normalized = normalized.split('://')[1]
    # 移除www.前缀
    if normalized.startswith('www.'):
        normalized = normalized[4:]
    # 移除路径部分(比如test.com/path只保留test.com)
    normalized = normalized.split('/')[0]
    return normalized

第二步:标准化目标URL列表

把从文件读取的fileURLs全部标准化,存到集合里(集合查找效率远高于列表):

fileURLs = urlReader()
print("原始待匹配URL列表:", fileURLs)
# 标准化所有待匹配URL
normalized_targets = {normalize_url(url) for url in fileURLs}
print("标准化后的待匹配URL:", normalized_targets)

第三步:遍历匹配数据

现在遍历你的data,把每个URL也标准化后,和目标集合对比:

for item in data:
    for url in item['urls']:
        normalized_current = normalize_url(url)
        # 检查标准化后的URL是否在目标集合中
        if normalized_current in normalized_targets:
            # 输出你需要的信息,比如id、configuredName和原始URL
            print(item['id'], item.get('configuredName', ''), url)

额外需求:直接匹配关键词

如果你不需要匹配完整域名,而是直接搜索test、youtube、microsoft这类关键词,可以把匹配逻辑改成这样:

target_keywords = {'test', 'youtube', 'microsoft'}
for item in data:
    for url in item['urls']:
        normalized_current = normalize_url(url)
        # 检查URL是否包含任意目标关键词
        if any(keyword in normalized_current for keyword in target_keywords):
            print(item['id'], item.get('configuredName', ''), url)

为什么这个方案可行?

不管原始URL是www.test.com、https://TEST.COM、test.com/path还是youtube.com,经过标准化后都会变成test.com或youtube.com——完全消除了格式差异的影响。不管你fileURLs里存的是带www的还是不带的,标准化后都会统一,不会再出现两次尝试里结果不一致的问题。

内容的提问来源于stack exchange,提问作者Hiddenx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:18:13