如何修改Stanford NLP代码生成水平格式的标注结果?
解决Stanford CoreNLP生成水平格式词形还原结果的方法
Stanford CoreNLP的命令行工具没有直接输出水平句式词形还原结果的选项,要实现你需要的格式,得用它的Java API写个简单的自定义程序,步骤如下:
1. 编写自定义Java代码
创建一个名为LemmaFormatter.java的文件,内容如下:
import edu.stanford.nlp.pipeline.*; import edu.stanford.nlp.ling.*; import java.io.*; import java.util.*; public class LemmaFormatter { public static void main(String[] args) throws IOException { // 配置NLP管道,启用需要的标注器 Properties props = new Properties(); props.setProperty("annotators", "tokenize,ssplit,pos,lemma"); StanfordCoreNLP pipeline = new StanfordCoreNLP(props); // 读取输入文件内容 File inputFile = new File("1.txt"); String content = new String(java.nio.file.Files.readAllBytes(inputFile.toPath())); // 初始化标注任务 Annotation document = new Annotation(content); pipeline.annotate(document); // 处理每个句子,拼接词形还原结果 List<CoreMap> sentences = document.get(CoreAnnotations.SentencesAnnotation.class); StringBuilder formattedResult = new StringBuilder(); for (CoreMap sentence : sentences) { for (CoreLabel token : sentence.get(CoreAnnotations.TokensAnnotation.class)) { // 保留原文本的前置空格/符号,保证格式和原文一致 formattedResult.append(token.get(CoreAnnotations.BeforeAnnotation.class)); // 追加词形还原结果 formattedResult.append(token.get(CoreAnnotations.LemmaAnnotation.class)); } } // 输出结果到控制台 System.out.println(formattedResult.toString()); // 将结果写入文件(可选) try (PrintWriter out = new PrintWriter("lemma_output.txt")) { out.println(formattedResult.toString()); } } }
2. 编译并运行程序
确保你的工作目录下有Stanford CoreNLP的jar包和1.txt文件,执行以下命令:
- 编译代码:
javac -cp "*" LemmaFormatter.java - 运行程序:
java -cp "*" LemmaFormatter
运行后,控制台会输出和原文句式一致的词形还原结果,同时也会生成lemma_output.txt文件保存结果。
关键说明
代码中通过BeforeAnnotation保留了原文本的空格、标点位置,还原后的句子结构和原文完全匹配,只是把每个词替换为对应的词形还原形式。
内容的提问来源于stack exchange,提问作者Bryan Jia
相关产品推荐
相关产品推荐

