如何用Python从URL列表中移除重复域名与无效HTML文件?
如何过滤URL列表:移除非HTML链接并去重域名
需求说明
给定一个URL列表,需要完成两项处理:
- 移除指向非HTML资源的链接(如PDF、JPG等文件)
- 同一域名仅保留根域名链接,剔除同域名下的其他子页面链接
示例输入
example_list = [ 'https://ocp.dc.gov/sites/default/files/dc/sites/ocp/publication/attachments/Report-of-Contracting-Activity-Part-I.pdf', 'https://the1955club.com/', 'https://the1955club.com/aboutus' ]
示例输出
new_list = ['https://the1955club.com/']
实现方案
用Python的urllib.parse模块解析URL,分两步完成过滤和去重:
from urllib.parse import urlparse example_list = [ 'https://ocp.dc.gov/sites/default/files/dc/sites/ocp/publication/attachments/Report-of-Contracting-Activity-Part-I.pdf', 'https://the1955club.com/', 'https://the1955club.com/aboutus' ] # 定义需要排除的非HTML文件后缀,可按需补充 excluded_extensions = {'.pdf', '.jpg', '.jpeg', '.png', '.gif', '.doc', '.docx', '.xls', '.xlsx'} # 第一步:过滤非HTML链接 filtered_html_links = [] for url in example_list: parsed_url = urlparse(url) # 检查URL路径是否以排除后缀结尾 if not any(parsed_url.path.endswith(ext) for ext in excluded_extensions): filtered_html_links.append(url) # 第二步:按域名去重,保留根域名链接 domain_root_map = {} for url in filtered_html_links: parsed_url = urlparse(url) domain = parsed_url.netloc # 拼接根域名链接 root_domain_url = f"{parsed_url.scheme}://{domain}/" # 利用字典键的唯一性确保每个域名只存一次根链接 if domain not in domain_root_map: domain_root_map[domain] = root_domain_url # 转换为最终结果列表 new_list = list(domain_root_map.values()) print(new_list)
代码说明
- 过滤非HTML链接:解析URL路径部分,判断是否属于预设的非HTML文件后缀,筛选出符合要求的HTML类链接。
- 域名去重:借助字典的键唯一性,以域名为键存储对应的根域名链接,确保每个域名仅保留一条根链接。
内容的提问来源于stack exchange,提问作者Kelly Tang
相关产品推荐
相关产品推荐

