You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中使用udpipe包实现多词聚类以识别关键表达?

Identifying Multi-Word Key Expressions with udpipe + Clustering

Hey there! Let’s tackle this problem—you’re already working with udpipe for text mining, and you want to group multiple words to spot expressions like from dusk till dawn? Totally doable by building on the co-occurrence work from that tutorial. Here’s how to approach it:

1. Start with Co-Occurrence Data (Building on the Tutorial)

First, recall that the tutorial uses co-occurrence of nouns/adjectives to build a network. We can expand this to include all relevant word types (prepositions, verbs, etc.—since expressions like your example mix prepositions and nouns) and use community detection clustering to group words that frequently appear together.

Step 1: Prep Your udpipe-Annotated Data

Assuming you’ve already run udpipe_annotate() on your text and have a data frame x with columns like doc_id, sentence_id, token, lemma, and upos (universal part-of-speech tags):

# Filter to keep relevant POS tags (adjust based on your needs)
# We'll keep nouns, verbs, prepositions, and adjectives here
filtered_tokens <- x[x$upos %in% c("NOUN", "VERB", "ADP", "ADJ"), ]

# Generate co-occurrence counts for words within the same sentence
# Using lemmas to group inflected forms (e.g., "dusk" and "dusks" count as one)
cooc_data <- udpipe::cooccurrence(
  x = filtered_tokens,
  term = "lemma",
  group = c("doc_id", "sentence_id")  # Co-occur within the same sentence
)

2. Cluster with Community Detection (Network-Based Clustering)

We’ll use the igraph package to turn co-occurrence data into a network, then run community detection to group words that often appear together. This works because words in expressions like from dusk till dawn will have strong co-occurrence links.

library(igraph)

# Create a graph from co-occurrence data (only keep high-frequency pairs to reduce noise)
cooc_graph <- graph_from_data_frame(
  cooc_data[cooc_data$cooc > 5, ],  # Adjust threshold based on your dataset size
  directed = FALSE
)

# Run Louvain community detection (great for finding dense subgroups in networks)
word_communities <- cluster_louvain(cooc_graph)

# View which words belong to which community
head(membership(word_communities))

Visualize the Clustered Network (Like the Tutorial’s Graph)

You can plot the network with communities colored differently to see your groups clearly—just like the tutorial’s visualization, but now with clustered word groups:

plot(
  word_communities, cooc_graph,
  vertex.size = degree(cooc_graph) * 2,  # Size nodes by how often they co-occur
  vertex.label.color = "black",
  edge.width = E(cooc_graph)$cooc / 10,  # Thicken edges for stronger co-occurrences
  main = "Clustered Word Co-Occurrence Network"
)

3. Bonus: N-Gram + TF-IDF Clustering for Exact Multi-Word Phrases

If you want to target fixed-length phrases (like 3-word expressions), you can extract n-grams first, then cluster them using TF-IDF to group semantically similar phrases:

# Extract 3-grams from your tokenized text
tri_grams <- udpipe::txt_ngrams(
  x = x$token,
  n = 3,
  skip = 0,
  collapse = " "
)

# Convert n-grams to a TF-IDF matrix to capture phrase importance
library(tm)
ngram_corpus <- VCorpus(VectorSource(tri_grams))
tfidf_dtm <- DocumentTermMatrix(
  ngram_corpus,
  control = list(weighting = weightTfIdf)
)

# Cluster with K-Means (adjust centers based on how many phrase groups you expect)
kmeans_clusters <- kmeans(as.matrix(tfidf_dtm), centers = 10)

# Inspect top phrases in each cluster
for (cluster_num in 1:10) {
  cat("\nCluster", cluster_num, "Top Phrases:\n")
  print(head(tri_grams[kmeans_clusters$cluster == cluster_num]))
}

Wrap-Up

Either approach will help you identify those multi-word key expressions:

  • Network community detection is great for finding flexible groups of words that often appear together (even in varying orders).
  • N-gram + TF-IDF clustering targets fixed-length phrases and groups semantically similar ones.

Play around with POS filters, co-occurrence thresholds, and cluster parameters to fit your specific dataset!

内容的提问来源于stack exchange,提问作者MysteryGuy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:36:58