使用mmap加速大文件搜索时报错:'mmap.mmap'对象无split属性
问题解决:AttributeError: 'mmap.mmap' object has no attribute 'split'
错误原因
mmap.mmap对象并非字符串或字节串类型,它没有内置的split()方法,直接对其调用split()才触发了这个错误。
修复方案
核心是先将mmap对象的内容读取为字节串(bytes),再执行split()操作;同时建议改用二进制模式打开文件,避免文本模式下的编码、换行符自动转换干扰mmap的正常工作。
修改后的代码
import re import mmap # 改用二进制模式打开文件,适配mmap的操作逻辑 with open('words.txt', 'rb') as f1, open('source.txt', 'rb') as f2, open('output.txt', 'w', encoding="utf8") as output_file: # 将文件映射到内存 file1_contents = mmap.mmap(f1.fileno(), 0, access=mmap.ACCESS_READ) file2_contents = mmap.mmap(f2.fileno(), 0, access=mmap.ACCESS_READ) # 先读取mmap内容为字节串,再执行split target_words = file2_contents.read().split() # 构建正则匹配模式 pattern = re.compile(b'(' + b'|'.join(target_words) + b')') # 逐行搜索匹配内容 for line in iter(file1_contents.readline(), b""): if pattern.search(line): # 将字节串解码为字符串后写入输出文件 output_file.write(line.decode('utf8'))
关键改动说明
- 文件打开模式:将
words.txt和source.txt改为二进制模式(rb),mmap对二进制数据的处理更稳定,也避免了文本模式下隐式的编码转换可能带来的问题。 - 读取mmap内容:通过
file2_contents.read()将mmap对象的内容读取为字节串,此时就能正常调用split()方法分割单词。 - 编码处理:写入输出文件时,将匹配到的字节行通过
decode('utf8')转换为UTF-8编码的字符串,适配输出文件的文本模式。
内容的提问来源于stack exchange,提问作者Bobby TB
相关产品推荐
相关产品推荐

