Python中忽略URL前缀后缀,匹配关键词并输出完整URL
解决URL格式不一致的匹配问题
你的两次尝试都卡在了URL格式差异上——要么只能精确匹配完全一致的字符串,要么匹配结果跟着fileURLs的格式变来变去。核心问题是没消除www.、http://、https://这些前缀的干扰,我们需要先把所有URL标准化处理,再进行匹配。
第一步:编写URL标准化函数
这个函数会把各种格式的URL转换成统一形式,去掉无关前缀、统一小写:
def normalize_url(url): # 转小写,避免大小写差异 normalized = str(url).lower() # 移除http://或https://前缀 if normalized.startswith(('http://', 'https://')): normalized = normalized.split('://')[1] # 移除www.前缀 if normalized.startswith('www.'): normalized = normalized[4:] # 移除路径部分(比如test.com/path只保留test.com) normalized = normalized.split('/')[0] return normalized
第二步:标准化目标URL列表
把从文件读取的fileURLs全部标准化,存到集合里(集合查找效率远高于列表):
fileURLs = urlReader() print("原始待匹配URL列表:", fileURLs) # 标准化所有待匹配URL normalized_targets = {normalize_url(url) for url in fileURLs} print("标准化后的待匹配URL:", normalized_targets)
第三步:遍历匹配数据
现在遍历你的data,把每个URL也标准化后,和目标集合对比:
for item in data: for url in item['urls']: normalized_current = normalize_url(url) # 检查标准化后的URL是否在目标集合中 if normalized_current in normalized_targets: # 输出你需要的信息,比如id、configuredName和原始URL print(item['id'], item.get('configuredName', ''), url)
额外需求:直接匹配关键词
如果你不需要匹配完整域名,而是直接搜索test、youtube、microsoft这类关键词,可以把匹配逻辑改成这样:
target_keywords = {'test', 'youtube', 'microsoft'} for item in data: for url in item['urls']: normalized_current = normalize_url(url) # 检查URL是否包含任意目标关键词 if any(keyword in normalized_current for keyword in target_keywords): print(item['id'], item.get('configuredName', ''), url)
为什么这个方案可行?
不管原始URL是www.test.com、https://TEST.COM、test.com/path还是youtube.com,经过标准化后都会变成test.com或youtube.com——完全消除了格式差异的影响。不管你fileURLs里存的是带www的还是不带的,标准化后都会统一,不会再出现两次尝试里结果不一致的问题。
内容的提问来源于stack exchange,提问作者Hiddenx
相关产品推荐
相关产品推荐

