You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行语义文本映射代码时遇np.nan无效文档错误求助

解决TF-IDF文本映射时的NaN值错误

错误原因

报错ValueError: np.nan is an invalid document, expected byte or unicode string的核心原因是:你的Excel数据中存在空值(NaN),而TfidfVectorizer仅能处理字符串类型的文本,无法识别非字符串的空值数据。

解决方案

在读取数据后,先对文本列的空值进行处理,提供两种可选方案:

方案1:将空值替换为空字符串

适合需要保留所有行数据的场景,把NaN转换为TF-IDF可识别的空字符串:

import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

# 读取Excel文件
df1 = pd.read_excel("E:/file1.xlsx")
df2 = pd.read_excel("E:/file2.xlsx")

# 处理空值:将NaN替换为空字符串并转为字符串类型
df1['Business descriptions'] = df1['Business descriptions'].fillna('').astype(str)
df2['Business licences'] = df2['Business licences'].fillna('').astype(str)

# 初始化TF-IDF向量器
tfidf_vectorizer = TfidfVectorizer()

# 合并文本数据
combined_text = list(df1['Business descriptions']) + list(df2['Business licences'])

# 生成TF-IDF矩阵
tfidf_matrix = tfidf_vectorizer.fit_transform(combined_text)

# 计算余弦相似度
cosine_sim = cosine_similarity(tfidf_matrix, tfidf_matrix)

# 定义描述和许可证的索引范围
desc_indices = range(len(df1))
lic_indices = range(len(df1), len(df1) + len(df2))

# 构建映射字典
mapping_dict = {}
for desc_idx in desc_indices:
    best_match_idx = max(lic_indices, key=lambda x: cosine_sim[desc_idx][x])
    mapping_dict[df1.loc[desc_idx, 'Business codes']] = df2.loc[best_match_idx - len(df1), 'Business licences']

# 生成结果DataFrame并保存
mapped_df = pd.DataFrame(list(mapping_dict.items()), columns=['Business codes', 'Mapped Business licences'])
mapped_df.to_excel('mapped_data.xlsx', index=False)

方案2:删除含空值的行

适合空值行无业务价值的场景,直接移除包含空值的记录:

import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

# 读取Excel文件
df1 = pd.read_excel("E:/file1.xlsx")
df2 = pd.read_excel("E:/file2.xlsx")

# 处理空值:删除包含空值的行
df1 = df1.dropna(subset=['Business descriptions'])
df2 = df2.dropna(subset=['Business licences'])

# 后续代码与方案1一致
tfidf_vectorizer = TfidfVectorizer()
combined_text = list(df1['Business descriptions']) + list(df2['Business licences'])
tfidf_matrix = tfidf_vectorizer.fit_transform(combined_text)
cosine_sim = cosine_similarity(tfidf_matrix, tfidf_matrix)

desc_indices = range(len(df1))
lic_indices = range(len(df1), len(df1) + len(df2))

mapping_dict = {}
for desc_idx in desc_indices:
    best_match_idx = max(lic_indices, key=lambda x: cosine_sim[desc_idx][x])
    mapping_dict[df1.loc[desc_idx, 'Business codes']] = df2.loc[best_match_idx - len(df1), 'Business licences']

mapped_df = pd.DataFrame(list(mapping_dict.items()), columns=['Business codes', 'Mapped Business licences'])
mapped_df.to_excel('mapped_data.xlsx', index=False)

说明

两种方案可根据你的数据情况选择:如果需要保留所有业务代码记录,优先选方案1;如果空值行没有分析意义,方案2能减少无效数据干扰。

内容的提问来源于stack exchange,提问作者mobinhb

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 09:13:09