You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在C#中去除单词后缀并还原动词原形?

动词原形还原实现方案

需求说明

需要将获取到的单词去除时态/后缀,还原为动词原形,例如:

  • playing → play
  • watches → watch
  • stopped → stop

尝试过Humanizer和OpenNlp工具,但不清楚具体用法,也没找到合适的实现方式,自己写的代码存在问题。

现有代码问题分析

用户提供的代码:

public List<string> changeWord(List<string> wordss,string baseUrl)
{
    string[] wordEnd = {"ing","es", "ies"};
    List<string> tags = getH1AndTitleTags(baseUrl);
    foreach(string tag in tags)
    {
        if (tag.Contains(wordEnd[0]))
        {
            tag.Replace("ing", "");
            tags.Add(tag);
        }
    }
    return tags;
}

存在以下问题:

  • 字符串修改无效:C#中string是不可变类型,tag.Replace()会返回新字符串,原tag不会被修改,代码中未接收返回值
  • 集合遍历异常:遍历tags时直接向集合中添加元素,会导致循环逻辑混乱,甚至抛出异常
  • 处理范围有限:只处理了"ing"后缀,未覆盖"es"、"ies"以及像"stopped"这类需要去双写尾字母的情况
  • 冗余参数:方法参数wordss未被使用,属于无效参数

工具的正确用法

1. Humanizer(推荐,简单易用)

首先安装Humanizer NuGet包,然后通过内置的动词变形方法直接还原原形:

using Humanizer;

// 单个单词还原示例
string playingBase = "playing".Verb().InflectTo(VerbForm.Infinitive); // 返回"play"
string watchesBase = "watches".Verb().InflectTo(VerbForm.Infinitive); // 返回"watch"
string stoppedBase = "stopped".Verb().InflectTo(VerbForm.Infinitive); // 返回"stop"

2. OpenNlp(适合NLP场景,需配合词性标注)

OpenNlp需要下载动词词形还原模型(如en-verb.bin),加载模型后结合词性标注实现还原:

using opennlp.tools.lemmatizer;
using opennlp.tools.util;

// 加载模型
LemmatizerModel model = new LemmatizerModel(new File("en-verb.bin"));
EnglishLemmatizer lemmatizer = new EnglishLemmatizer(model);

// 需指定单词词性,VBG为动名词、VBZ为第三人称单数、VBD为过去式
string playingBase = lemmatizer.Lemmatize("playing", "VBG"); // 返回"play"
string watchesBase = lemmatizer.Lemmatize("watches", "VBZ"); // 返回"watch"
string stoppedBase = lemmatizer.Lemmatize("stopped", "VBD"); // 返回"stop"

注意:OpenNlp需要先对单词做词性标注,才能更准确地还原原形,可配合OpenNlp的词性标注器使用。

改进后的代码示例

基于Humanizer实现的完整代码,修正原逻辑问题:

using Humanizer;
using System.Collections.Generic;
using System.Linq;

public List<string> GetVerbBaseForms(string baseUrl)
{
    // 获取标签内容
    List<string> tagContents = getH1AndTitleTags(baseUrl);
    List<string> baseFormVerbs = new List<string>();

    foreach (string content in tagContents)
    {
        // 拆分内容为单个单词(处理空格、标点分隔的情况)
        string[] words = content.Split(new[] {' ', '.', ',', '!'}, StringSplitOptions.RemoveEmptyEntries);
        
        foreach (string word in words)
        {
            // 还原为动词原形
            string baseForm = word.Verb().InflectTo(VerbForm.Infinitive);
            baseFormVerbs.Add(baseForm);
        }
    }

    return baseFormVerbs;
}

内容的提问来源于stack exchange,提问作者Mr_Programmer002

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 04:31:00