R语言代词分析:如何获取分词后句子中代词的人称与类型信息?
Great question! When you're working with tokenized text like your example (sent <- "She likes long walks on the beach with her dogs."), there are several solid NLP tools and libraries that can pull pronoun-specific details like person (1st/2nd/3rd) and type (possessive, personal, etc.). Below are the most practical options, tailored for both R and Python since your code snippet uses R:
Python Libraries
spaCy
This is one of the most user-friendly and powerful options for this task. It comes with pre-trained models that directly expose morphological features like pronoun person and possessive status. Here's a quick example:import spacy # Load a pre-trained English model nlp = spacy.load("en_core_web_sm") doc = nlp("She likes long walks on the beach with her dogs.") # Iterate through tokens and filter pronouns for token in doc: if token.pos_ == "PRON": person = token.morph.get("Person") is_possessive = token.morph.get("Possessive") print(f"Token: {token.text} | Person: {person} | Possessive: {is_possessive}")Output will show
Sheas 3rd person, non-possessive, andheras 3rd person possessive.NLTK
A classic, beginner-friendly library. It uses Penn Treebank POS tags to identify pronoun types (e.g.,PRPfor personal pronouns,PRP$for possessive pronouns), and you can add simple logic to map tokens to their person category:import nltk from nltk.tag import pos_tag nltk.download("punkt") nltk.download("averaged_perceptron_tagger") tokens = nltk.word_tokenize("She likes long walks on the beach with her dogs.") tagged_tokens = pos_tag(tokens) for word, tag in tagged_tokens: if tag in ["PRP", "PRP$"]: # Map token to person if word.lower() in ["i", "me", "my", "mine"]: person = "1st" elif word.lower() in ["you", "your", "yours"]: person = "2nd" else: person = "3rd" # Determine pronoun type pron_type = "Possessive" if tag == "PRP$" else "Personal" print(f"Token: {word} | Person: {person} | Type: {pron_type}")
R Libraries
Since your original code uses R, here are two great options tailored to the ecosystem:
spaCyR
The R wrapper for spaCy, giving you access to all the same morphological features as the Python version. Here's how to use it:library(spacyr) # Initialize spaCy with an English model spacy_initialize(model = "en_core_web_sm") # Parse the text and extract relevant pronoun features parsed_doc <- spacy_parse( "She likes long walks on the beach with her dogs.", tag = TRUE, morph = TRUE ) # Filter for pronouns and select key columns subset(parsed_doc, pos == "PRON")[, c("token", "morph_person", "morph_possessive")]This will return a dataframe with the pronoun text, its person category, and whether it's possessive.
udpipe
A library built around Universal Dependencies (UD) annotations, which standardize morphological features across languages. It makes it easy to pull pronoun-specific attributes:library(udpipe) # Download and load a pre-trained English UD model ud_model <- udpipe_download_model(language = "english") ud_model_obj <- udpipe_load_model(ud_model$file_model) # Annotate the text annotated_text <- udpipe_annotate(ud_model_obj, "She likes long walks on the beach with her dogs.") annotated_df <- as.data.frame(annotated_text) # Filter pronouns and extract person/possessive features subset(annotated_df, upos == "PRON")[, c("token", "feats_Person", "feats_Poss")]The
feats_Personcolumn will show1,2, or3for the pronoun's person, andfeats_Posswill mark whether it's possessive.
For most use cases, spaCy (either Python or R) is the best bet—it’s intuitive, requires minimal custom code, and handles edge cases (like reflexive pronouns) out of the box.
内容的提问来源于stack exchange,提问作者Christopher Costello

