如何在Java中提取文本中的各类名词(含Common noun、Proper noun)
Hey there! Extracting common and proper nouns from text in Java can be done in a couple of practical ways—from quick rule-based checks (great for basic proper noun extraction) to using robust NLP libraries that handle both noun types accurately. Let’s break down each approach:
1. Rule-Based Approach (Quick Proper Noun Extraction)
Proper nouns are typically capitalized (with exceptions like "iPhone"), so we can leverage that pattern. We just need to avoid mistaking words at the start of sentences for proper nouns. Here’s a simple implementation:
import java.util.ArrayList; import java.util.List; import java.util.regex.Matcher; import java.util.regex.Pattern; public class RuleBasedNounExtractor { public static List<String> extractProperNouns(String text) { List<String> properNouns = new ArrayList<>(); // Split text into individual sentences String[] sentences = text.split("[.!?]+\\s*"); Pattern capitalizedWordPattern = Pattern.compile("\\b[A-Z][a-z]*\\b"); for (String sentence : sentences) { Matcher matcher = capitalizedWordPattern.matcher(sentence); boolean isFirstWord = true; while (matcher.find()) { String word = matcher.group(); if (isFirstWord) { isFirstWord = false; continue; // Skip the first word of each sentence } properNouns.add(word); } } return properNouns; } public static void main(String[] args) { String text = "Steven lives in London. He visited Africa last Monday. The bridge over the river is old."; System.out.println("Proper Nouns: " + extractProperNouns(text)); } }
Limitations: This works for basic cases but fails with hyphenated proper nouns ("New York") or lowercase proper nouns. Rule-based methods aren’t reliable for common nouns since they don’t have consistent formatting cues.
2. Using NLP Libraries (Accurate for Both Noun Types)
For reliable extraction of both common and proper nouns, use a Part-of-Speech (POS) tagging library. Two popular options are Stanford CoreNLP and OpenNLP.
Option A: Stanford CoreNLP
Stanford CoreNLP provides state-of-the-art POS tagging. First, add these Maven dependencies to your project:
<dependency> <groupId>edu.stanford.nlp</groupId> <artifactId>stanford-corenlp</artifactId> <version>4.5.4</version> </dependency> <dependency> <groupId>edu.stanford.nlp</groupId> <artifactId>stanford-corenlp</artifactId> <version>4.5.4</version> <classifier>models</classifier> </dependency>
Then use this code to extract both noun types:
import edu.stanford.nlp.ling.CoreAnnotations; import edu.stanford.nlp.ling.CoreLabel; import edu.stanford.nlp.pipeline.Annotation; import edu.stanford.nlp.pipeline.StanfordCoreNLP; import edu.stanford.nlp.util.CoreMap; import java.util.ArrayList; import java.util.List; import java.util.Properties; public class StanfordNlpNounExtractor { public static void main(String[] args) { Properties props = new Properties(); props.setProperty("annotators", "tokenize,ssplit,pos"); StanfordCoreNLP pipeline = new StanfordCoreNLP(props); String text = "Steven lives in London. He visited Africa last Monday. The bridge over the river is old. Happiness is important."; Annotation document = new Annotation(text); pipeline.annotate(document); List<String> commonNouns = new ArrayList<>(); List<String> properNouns = new ArrayList<>(); for (CoreMap sentence : document.get(CoreAnnotations.SentencesAnnotation.class)) { for (CoreLabel token : sentence.get(CoreAnnotations.TokensAnnotation.class)) { String posTag = token.get(CoreAnnotations.PartOfSpeechAnnotation.class); String word = token.get(CoreAnnotations.TextAnnotation.class); // Common noun tags: NN (singular), NNS (plural) if (posTag.equals("NN") || posTag.equals("NNS")) { commonNouns.add(word); } // Proper noun tags: NNP (singular), NNPS (plural) else if (posTag.equals("NNP") || posTag.equals("NNPS")) { properNouns.add(word); } } } System.out.println("Common Nouns: " + commonNouns); System.out.println("Proper Nouns: " + properNouns); } }
Sample Output:
Common Nouns: [bridge, river, Happiness]
Proper Nouns: [Steven, London, Africa, Monday]
Option B: OpenNLP
OpenNLP is a lightweight alternative. Add this Maven dependency:
<dependency> <groupId>org.apache.opennlp</groupId> <artifactId>opennlp-tools</artifactId> <version>2.3.2</version> </dependency>
Download the English POS tagger model (en-pos-maxent.bin) from the OpenNLP website and place it in your project’s resources folder. Then use this code:
import opennlp.tools.postag.POSModel; import opennlp.tools.postag.POSTaggerME; import opennlp.tools.tokenize.SimpleTokenizer; import java.io.FileInputStream; import java.io.IOException; import java.io.InputStream; import java.util.ArrayList; import java.util.List; public class OpenNlpNounExtractor { public static void main(String[] args) throws IOException { // Load the POS model InputStream modelIn = new FileInputStream("src/main/resources/en-pos-maxent.bin"); POSModel model = new POSModel(modelIn); POSTaggerME tagger = new POSTaggerME(model); String text = "Steven lives in London. He visited Africa last Monday. The bridge over the river is old. Happiness is important."; String[] tokens = SimpleTokenizer.INSTANCE.tokenize(text); String[] tags = tagger.tag(tokens); List<String> commonNouns = new ArrayList<>(); List<String> properNouns = new ArrayList<>(); for (int i = 0; i < tokens.length; i++) { String tag = tags[i]; String word = tokens[i]; if (tag.startsWith("NN") && !tag.startsWith("NNP")) { // Common nouns commonNouns.add(word); } else if (tag.startsWith("NNP")) { // Proper nouns properNouns.add(word); } } System.out.println("Common Nouns: " + commonNouns); System.out.println("Proper Nouns: " + properNouns); modelIn.close(); } }
Key Takeaways
- Rule-based methods are fast but error-prone; use them only for simple proper noun extraction.
- NLP libraries are far more accurate, especially for common nouns and edge cases (like compound nouns or lowercase proper nouns).
- Always use the latest library versions to get the best performance and accuracy.
内容的提问来源于stack exchange,提问作者nixxo_raa

