如何用Python按百分比对比用户输入与现有文件内容?
实现用户输入与文件内容的相似度百分比对比
需求是通过Python计算用户输入内容与指定文件内容的单词相似度百分比,示例:
- 目标文件内容:
this is a butifull day(共5个单词) - 用户输入内容:
what a butifull day right?(包含目标文件中的3个单词) - 计算得相似度约为60%(3/5×100)
需要修改原有代码中elif user in client:对应的判断逻辑,原代码如下:
def cindex(msg): #add client input to "client" file with open("client_index.txt", "a") as client: client.write(msg) client.write("\n") client.close() #open and compare the files with open("user_index.txt", "r") as user: with open("client_index.txt", "r") as client: same = set(user).intersection(client) for line in same: print(line, end='') if msg == "[Please reactivate the user!]": driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys('same') elif user in client: driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys('yap') else: driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys('nope')
修改后的代码
import re def cindex(msg): # 将用户输入追加到client文件 with open("client_index.txt", "a") as client: client.write(msg) client.write("\n") # 读取目标文件和用户输入历史文件内容 with open("user_index.txt", "r") as user_file: user_content = user_file.read().strip() with open("client_index.txt", "r") as client_file: # 取最后一行即当前用户输入 client_lines = client_file.readlines() current_input = client_lines[-1].strip() if client_lines else "" # 提取单词(忽略标点,转为小写统一格式) def extract_words(text): # 用正则匹配单词,去掉标点,转为小写 return set(re.findall(r'\b\w+\b', text.lower())) user_words = extract_words(user_content) input_words = extract_words(current_input) if not user_words: # 目标文件为空时直接返回无相似度 similarity = 0 else: # 计算共同单词数量 common_words = user_words.intersection(input_words) # 计算相似度百分比 similarity = (len(common_words) / len(user_words)) * 100 # 原逻辑保留,替换elif的判断为基于相似度的逻辑 if msg == "[Please reactivate the user!]": driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys('same') elif similarity >= 50: # 可根据需求调整阈值 driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys(f'yap(相似度{round(similarity, 1)}%)') else: driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys(f'nope(相似度{round(similarity, 1)}%)')
关键修改说明
- 新增
extract_words函数:用正则提取文本中的单词,统一转为小写并去除标点,避免因大小写或标点导致的匹配误差 - 读取文件内容改为提取完整文本,而非按行处理,确保能正确统计所有单词
- 计算相似度:以目标文件(
user_index.txt)的单词总数为分母,共同单词数为分子,计算百分比 - 修改
elif的判断逻辑:可自定义相似度阈值(示例设为50%),超过阈值时发送包含相似度的提示,否则发送低相似度提示 - 移除了不必要的
client.close(),因为with语句会自动关闭文件
内容的提问来源于stack exchange,提问作者Elior
相关产品推荐
相关产品推荐

