You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Django中实现基于标题匹配3-4个相同词的相关新闻聚合

实现基于标题词汇匹配的相关新闻推荐(3-4个相同词汇)

Hey William, 针对你在新闻详情页需要匹配3个或4个相同词汇的相关新闻需求,我结合你的Django News 模型整理了几个实用的实现方案,涵盖中小数据量和大数据量场景:

方案一:基础字符串处理(适合中小站点)

这个方案逻辑简单,不需要额外的数据库扩展,直接用Python字符串处理+Django ORM就能实现:

步骤说明

  • 先清洗当前新闻的标题:拆分词汇、过滤无意义的停用词(比如英文的a/the、中文的“的/了”)、去重
  • 遍历其他新闻,计算与当前标题的共同词汇数量
  • 筛选出共同词汇数在3-4之间的新闻,再按匹配数量排序

代码实现

from django.db.models import Q
import re
# 如果处理英文,用nltk停用词;中文的话替换成中文停用词表
from nltk.corpus import stopwords
# 第一次使用需要下载停用词:nltk.download('stopwords')

def get_related_news(current_news):
    # 1. 清洗当前新闻标题,提取有效词汇
    current_title = current_news.title.lower()
    # 用正则拆分词汇(英文适用),中文请替换为jieba分词
    current_words = re.findall(r'\b\w+\b', current_title)
    # 过滤停用词,避免无意义词汇干扰匹配
    stop_words = set(stopwords.words('english'))
    filtered_words = [word for word in current_words if word not in stop_words]
    # 去重,避免重复词汇影响计数
    unique_valid_words = list(set(filtered_words))
    
    # 如果当前标题有效词汇不足3个,直接返回空列表(可根据需求调整)
    if len(unique_valid_words) < 3:
        return News.objects.none()
    
    # 2. 筛选符合条件的相关新闻
    related_news_list = []
    # 排除当前新闻本身
    candidate_news = News.objects.exclude(id=current_news.id)
    
    for news in candidate_news:
        # 清洗候选新闻标题
        news_title = news.title.lower()
        news_words = re.findall(r'\b\w+\b', news_title)
        news_filtered = [word for word in news_words if word not in stop_words]
        news_valid_set = set(news_filtered)
        
        # 计算共同词汇数量
        common_word_count = len(set(unique_valid_words) & news_valid_set)
        # 匹配3或4个相同词汇的条件
        if 3 <= common_word_count <= 4:
            related_news_list.append(news)
    
    # 按匹配词汇数降序排序,匹配4个的优先展示
    related_news_list.sort(
        key=lambda x: len(set(re.findall(r'\b\w+\b', x.title.lower()) & set(filtered_words))),
        reverse=True
    )
    
    return related_news_list

方案二:数据库优化方案(适合大数据量站点)

如果你的新闻数据量很大(比如上万条),遍历所有新闻会影响性能,这个方案先通过数据库筛选缩小候选范围,再计算匹配数:

代码实现

from django.db.models import Q
import re
from nltk.corpus import stopwords

def get_related_news(current_news):
    current_title = current_news.title.lower()
    current_words = re.findall(r'\b\w+\b', current_title)
    stop_words = set(stopwords.words('english'))
    filtered_words = [word for word in current_words if word not in stop_words]
    unique_valid_words = list(set(filtered_words))
    
    if len(unique_valid_words) < 3:
        return News.objects.none()
    
    # 第一步:用数据库Q条件筛选出包含至少一个有效词汇的新闻,缩小候选范围
    q_condition = Q()
    for word in unique_valid_words:
        q_condition |= Q(title__icontains=word)
    
    candidate_news = News.objects.exclude(id=current_news.id).filter(q_condition)
    
    # 第二步:在候选新闻中计算精确匹配数
    related_news_list = []
    for news in candidate_news:
        news_valid_set = set(re.findall(r'\b\w+\b', news.title.lower())) - stop_words
        common_count = len(news_valid_set & set(unique_valid_words))
        if 3 <= common_count <=4:
            related_news_list.append(news)
    
    # 排序后返回
    related_news_list.sort(key=lambda x: len(set(re.findall(r'\b\w+\b', x.title.lower()) & set(filtered_words))), reverse=True)
    return related_news_list

关键注意事项

  • 中文适配:如果是中文新闻,需要把正则拆分替换为中文分词库(比如jieba),停用词换成中文停用词表(比如哈工大停用词表)
  • 性能优化:数据量极大时,可以预计算每个新闻的有效词汇集合,存储在JSONField字段中,查询时直接对比集合,避免重复分词
  • 缓存优化:用Django的缓存框架(比如cache.set())把相关新闻结果缓存1-2小时,减少重复计算

内容的提问来源于stack exchange,提问作者William

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:55:04