You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python按百分比对比用户输入与现有文件内容?

实现用户输入与文件内容的相似度百分比对比

需求是通过Python计算用户输入内容与指定文件内容的单词相似度百分比,示例:

  • 目标文件内容:this is a butifull day(共5个单词)
  • 用户输入内容:what a butifull day right?(包含目标文件中的3个单词)
  • 计算得相似度约为60%(3/5×100)

需要修改原有代码中elif user in client:对应的判断逻辑,原代码如下:

def cindex(msg):
    #add client input to "client" file
    with open("client_index.txt", "a") as client:
        client.write(msg)
        client.write("\n")
        client.close()
    #open and compare the files
    with open("user_index.txt", "r") as user:
        with open("client_index.txt", "r") as client:
            same = set(user).intersection(client)
            for line in same:
                print(line, end='')
            if msg == "[Please reactivate the user!]":
                driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys('same')
            elif user in client:
                driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys('yap')
            else:
                driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys('nope')

修改后的代码

import re

def cindex(msg):
    # 将用户输入追加到client文件
    with open("client_index.txt", "a") as client:
        client.write(msg)
        client.write("\n")
    # 读取目标文件和用户输入历史文件内容
    with open("user_index.txt", "r") as user_file:
        user_content = user_file.read().strip()
    with open("client_index.txt", "r") as client_file:
        # 取最后一行即当前用户输入
        client_lines = client_file.readlines()
        current_input = client_lines[-1].strip() if client_lines else ""

    # 提取单词(忽略标点,转为小写统一格式)
    def extract_words(text):
        # 用正则匹配单词,去掉标点,转为小写
        return set(re.findall(r'\b\w+\b', text.lower()))
    
    user_words = extract_words(user_content)
    input_words = extract_words(current_input)

    if not user_words:
        # 目标文件为空时直接返回无相似度
        similarity = 0
    else:
        # 计算共同单词数量
        common_words = user_words.intersection(input_words)
        # 计算相似度百分比
        similarity = (len(common_words) / len(user_words)) * 100

    # 原逻辑保留,替换elif的判断为基于相似度的逻辑
    if msg == "[Please reactivate the user!]":
        driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys('same')
    elif similarity >= 50:  # 可根据需求调整阈值
        driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys(f'yap(相似度{round(similarity, 1)}%)')
    else:
        driver.find_element("xpath", "//textarea[@id='chat-windows-message-textarea']").send_keys(f'nope(相似度{round(similarity, 1)}%)')

关键修改说明

  • 新增extract_words函数:用正则提取文本中的单词,统一转为小写并去除标点,避免因大小写或标点导致的匹配误差
  • 读取文件内容改为提取完整文本,而非按行处理,确保能正确统计所有单词
  • 计算相似度:以目标文件(user_index.txt)的单词总数为分母,共同单词数为分子,计算百分比
  • 修改elif的判断逻辑:可自定义相似度阈值(示例设为50%),超过阈值时发送包含相似度的提示,否则发送低相似度提示
  • 移除了不必要的client.close(),因为with语句会自动关闭文件

内容的提问来源于stack exchange,提问作者Elior

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 06:35:37