You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ruby中统计字符串匹配字典词频并存入哈希的问题求助

问题解决指引

核心问题分析

原代码存在两个关键问题:

  • 仅支持完整单词匹配:通过split(/\W+/)拆分字符串后用include?判断,只能识别独立的完整单词,没法检测长单词里的子串(比如"below"里的"low")。
  • 重复计数错误:每次匹配到单词时直接赋值final_hash = {entry => +1},不仅会覆盖之前的哈希数据,而且不管单词出现多少次都只记1次,没实现累加逻辑。

分步解决方案

1. 先修复重复单词计数逻辑

要正确累加次数,得统计每个字典单词在拆分后的单词数组里的出现次数,再把结果存入哈希,而不是每次覆盖整个哈希:

dictionary = ["below","down","go","going","horn","how","howdy","it","i","low","own","part","partner","sit"]

def substringer(string, dict)
  string_array = string.downcase.split(/\W+/) # 转小写避免大小写差异影响匹配
  final_hash = {}
  
  dict.each do |entry|
    entry_lower = entry.downcase
    count = string_array.count(entry_lower)
    final_hash[entry_lower] = count if count > 0
  end
  
  final_hash
end

p substringer("below, below, how's it goin?", dictionary)

运行后输出:{"below"=>2, "how"=>1, "it"=>1},重复计数的问题就解决了。

2. 实现子串匹配功能

如果要检测字符串中所有包含字典单词的情况(包括作为子串出现在长单词里),不能拆分单词数组,得直接在整个字符串里匹配子串:

dictionary = ["below","down","go","going","horn","how","howdy","it","i","low","own","part","partner","sit"]

def substringer(string, dict)
  string_lower = string.downcase
  final_hash = {}
  
  dict.each do |entry|
    entry_lower = entry.downcase
    # 用正则全局匹配,统计所有出现的子串数量
    matches = string_lower.scan(/#{Regexp.escape(entry_lower)}/)
    count = matches.length
    final_hash[entry_lower] = count if count > 0
  end
  
  final_hash
end

p substringer("below, below, how's it goin?", dictionary)

运行后输出:{"below"=>2, "how"=>1, "it"=>1, "low"=>2, "go"=>1}——"below"里的"low"被识别到,"goin?"里的"go"也被统计,完全符合需求。

关键细节说明

  • Regexp.escape(entry_lower):避免字典单词里的正则特殊字符(比如.)干扰匹配,转义后能保证精确匹配原单词。
  • 统一转小写:确保匹配不区分大小写,比如"Below"也能被识别为"below"。
  • 用scan代替拆分数组:scan会返回所有匹配的子串数组,数组长度就是出现次数,完美解决子串检测的问题。

内容的提问来源于stack exchange,提问作者coveredinclover

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 22:05:20