You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用mmap加速大文件搜索时报错:'mmap.mmap'对象无split属性

问题解决:AttributeError: 'mmap.mmap' object has no attribute 'split'

错误原因

mmap.mmap对象并非字符串或字节串类型,它没有内置的split()方法,直接对其调用split()才触发了这个错误。

修复方案

核心是先将mmap对象的内容读取为字节串(bytes),再执行split()操作;同时建议改用二进制模式打开文件,避免文本模式下的编码、换行符自动转换干扰mmap的正常工作。

修改后的代码

import re
import mmap

# 改用二进制模式打开文件,适配mmap的操作逻辑
with open('words.txt', 'rb') as f1, open('source.txt', 'rb') as f2, open('output.txt', 'w', encoding="utf8") as output_file:

    # 将文件映射到内存
    file1_contents = mmap.mmap(f1.fileno(), 0, access=mmap.ACCESS_READ)
    file2_contents = mmap.mmap(f2.fileno(), 0, access=mmap.ACCESS_READ)

    # 先读取mmap内容为字节串,再执行split
    target_words = file2_contents.read().split()
    # 构建正则匹配模式
    pattern = re.compile(b'(' + b'|'.join(target_words) + b')')

    # 逐行搜索匹配内容
    for line in iter(file1_contents.readline(), b""):
        if pattern.search(line):
            # 将字节串解码为字符串后写入输出文件
            output_file.write(line.decode('utf8'))

关键改动说明

  • 文件打开模式:将words.txt和source.txt改为二进制模式(rb),mmap对二进制数据的处理更稳定,也避免了文本模式下隐式的编码转换可能带来的问题。
  • 读取mmap内容:通过file2_contents.read()将mmap对象的内容读取为字节串,此时就能正常调用split()方法分割单词。
  • 编码处理:写入输出文件时,将匹配到的字节行通过decode('utf8')转换为UTF-8编码的字符串,适配输出文件的文本模式。

内容的提问来源于stack exchange,提问作者Bobby TB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 02:47:40