You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大数据量下关键词与域名文件的高效匹配需求及代码优化

大数量级域名与关键词匹配的性能优化

需求说明

需在domains.txt文件中匹配keywords.txt内的关键词,两类文件每行分别对应一个域名和关键词;由于数据量极大(含2.8亿条域名数据),需重点考虑性能优化。

现有代码

def search_keyword_domain():
    # 仅需在域名文件中搜索关键词,无需修改原文件
    # 每行对应一个关键词和一个域名
    # 因数据量庞大,请重点考虑性能
    
    with open("result.txt", "a") as result:
        result.writelines(line)


def search_keyword():
    with open('domains.txt', 'r') as d:
        for line in d:
            line.strip()
        d.close()

    with open('keywords.txt', 'r') as f:
        for line in f:
            line = line.strip()
            search_keyword_domain(line)
        f.close()


if __name__ == '__main__':
    search_keyword()

示例文件

keywords.txt(共180个关键词)

google
messi
apple

domains.txt(共2.8亿条域名)

google.com
ronaldovsmess.com
anapple.com

性能优化方案

原代码存在逻辑缺失(未传递关键词到搜索函数、遍历域名未做有效处理),且完全未考虑大数量级数据的性能问题,以下是针对性优化方案:

  1. 关键词预处理:将所有关键词存入集合,集合的成员查询为O(1)时间复杂度,远快于列表遍历
  2. 单次IO绑定:避免重复打开/关闭文件,一次性加载关键词后,逐行处理域名并写入结果
  3. 低内存占用:不一次性加载所有域名到内存,利用Python文件对象的迭代器特性逐行读取处理
  4. 高效字符串匹配:直接使用底层C实现的in操作符做子串匹配,速度远快于自定义遍历逻辑

优化后的代码:

def match_keywords_to_domains():
    # 加载关键词到集合,预处理时过滤空行、统一大小写(可按需调整)
    with open('keywords.txt', 'r', encoding='utf-8') as f:
        keywords = {line.strip().lower() for line in f if line.strip()}
    
    # 逐行处理域名,匹配后写入结果
    with open('domains.txt', 'r', encoding='utf-8') as domains_file, \
         open('result.txt', 'w', encoding='utf-8') as result_file:
        for domain_line in domains_file:
            domain = domain_line.strip().lower()
            # 检查是否存在匹配的关键词
            if any(keyword in domain for keyword in keywords):
                result_file.write(domain_line)  # 保留原行格式,如需处理后写入可替换为domain + '\n'

if __name__ == '__main__':
    match_keywords_to_domains()

进阶优化建议

  • 多进程并行处理:利用multiprocessing模块将域名文件分块,多进程并行匹配,适合CPU密集型场景
  • 正则预编译:若需复杂匹配规则(如关键词不能是其他单词的子串),预编译正则表达式集合:
    import re
    keywords_re = re.compile('|'.join(re.escape(k) for k in keywords))
    # 匹配时使用:if keywords_re.search(domain)
    
  • 内存映射IO:使用mmap将文件映射到内存,减少磁盘IO开销
  • 无效行过滤:跳过空行或格式错误的行,减少不必要的计算

内容的提问来源于stack exchange,提问作者abhdhasvbfvf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 19:36:18