You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从URL列表中移除重复域名与无效HTML文件?

如何过滤URL列表:移除非HTML链接并去重域名

需求说明

给定一个URL列表,需要完成两项处理:

  1. 移除指向非HTML资源的链接(如PDF、JPG等文件)
  2. 同一域名仅保留根域名链接,剔除同域名下的其他子页面链接

示例输入

example_list = [
    'https://ocp.dc.gov/sites/default/files/dc/sites/ocp/publication/attachments/Report-of-Contracting-Activity-Part-I.pdf', 
    'https://the1955club.com/', 
    'https://the1955club.com/aboutus'
]

示例输出

new_list = ['https://the1955club.com/']

实现方案

用Python的urllib.parse模块解析URL,分两步完成过滤和去重:

from urllib.parse import urlparse

example_list = [
    'https://ocp.dc.gov/sites/default/files/dc/sites/ocp/publication/attachments/Report-of-Contracting-Activity-Part-I.pdf', 
    'https://the1955club.com/', 
    'https://the1955club.com/aboutus'
]

# 定义需要排除的非HTML文件后缀,可按需补充
excluded_extensions = {'.pdf', '.jpg', '.jpeg', '.png', '.gif', '.doc', '.docx', '.xls', '.xlsx'}

# 第一步:过滤非HTML链接
filtered_html_links = []
for url in example_list:
    parsed_url = urlparse(url)
    # 检查URL路径是否以排除后缀结尾
    if not any(parsed_url.path.endswith(ext) for ext in excluded_extensions):
        filtered_html_links.append(url)

# 第二步:按域名去重,保留根域名链接
domain_root_map = {}
for url in filtered_html_links:
    parsed_url = urlparse(url)
    domain = parsed_url.netloc
    # 拼接根域名链接
    root_domain_url = f"{parsed_url.scheme}://{domain}/"
    # 利用字典键的唯一性确保每个域名只存一次根链接
    if domain not in domain_root_map:
        domain_root_map[domain] = root_domain_url

# 转换为最终结果列表
new_list = list(domain_root_map.values())
print(new_list)

代码说明

  1. 过滤非HTML链接:解析URL路径部分,判断是否属于预设的非HTML文件后缀,筛选出符合要求的HTML类链接。
  2. 域名去重:借助字典的键唯一性,以域名为键存储对应的根域名链接,确保每个域名仅保留一条根链接。

内容的提问来源于stack exchange,提问作者Kelly Tang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 20:50:23