将gensim w2v转Tensorboard tsv遇TypeError:需bytes类对象而非str
解决Gensim Word2Vec转TensorBoard TSV时的TypeError问题
你现在肯定被文件读写模式和数据类型不匹配的问题搞晕了——一会儿要字符串,一会儿要字节,确实挺闹心的。我来帮你理清楚问题根源,再给你两种靠谱的解决思路:
核心问题分析
你之前为了解决write() argument must be str, not bytes,给open加了b(二进制模式),但这就要求所有写入文件的内容必须是bytes类型。而你后来用vector_row.encode('UTF-8')无效,大概率是因为vector_row本身不是字符串(比如是向量的数值列表/数组),直接调用encode当然会报错。
反过来,如果不用二进制模式(文本模式),那就要把所有bytes类型的内容(比如模型里的词可能是bytes格式)转成字符串,这样就能避免类型不匹配的问题。
方案1:用文本模式写入(推荐,因为TSV是文本文件)
这种方式下,我们全程处理字符串,避免字节和字符串的来回转换混乱:
from gensim.models import Word2Vec # 加载你的Word2Vec模型 model = Word2Vec.load("your_w2v_model.model") # 用文本模式打开两个TSV文件,指定编码为utf-8 with open('vectors.tsv', 'w+', encoding='utf-8') as vec_file, \ open('metadata.tsv', 'w+', encoding='utf-8') as meta_file: # 遍历模型中的每个词 for word in model.wv.index_to_key: # 处理词:如果词是bytes类型,先解码成字符串 if isinstance(word, bytes): word_str = word.decode('utf-8') else: word_str = str(word) # 写入元数据文件(每个词占一行) meta_file.write(f"{word_str}\n") # 把向量转换成制表符分隔的字符串,再写入向量文件 vector_str = '\t'.join(map(str, model.wv[word])) + '\n' vec_file.write(vector_str)
方案2:用二进制模式写入(严格处理字节)
如果你坚持用二进制模式,那要确保每一项写入的内容都是bytes:
from gensim.models import Word2Vec model = Word2Vec.load("your_w2v_model.model") # 二进制模式打开文件 with open('vectors.tsv', 'wb+') as vec_file, \ open('metadata.tsv', 'wb+') as meta_file: for word in model.wv.index_to_key: # 处理词:字符串转字节,字节直接用 if isinstance(word, str): word_bytes = word.encode('utf-8') else: word_bytes = word meta_file.write(word_bytes + b'\n') # 换行符也要用字节格式 # 处理向量:先转成字符串,再编码成字节 vector_str = '\t'.join(map(str, model.wv[word])) + '\n' vector_bytes = vector_str.encode('utf-8') vec_file.write(vector_bytes)
为什么你之前的encode无效?
你之前尝试的vector_row = vector_row.encode('UTF-8'),如果vector_row是向量的数值数组/列表,那它根本不是字符串,自然没法调用encode方法。必须先把向量转换成制表符分隔的字符串,再去编码成字节,这才是正确的顺序。
内容的提问来源于stack exchange,提问作者OverflowingTheGlass
相关产品推荐
相关产品推荐

