You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中基于文本DataFrame构建无向概念图?

Build an Undirected Concept Graph from Your Text DataFrame

Got it, let's walk through how to create that undirected concept graph using your 100-sample text DataFrame. I’ve tackled similar tasks before, so here’s a practical, step-by-step approach that balances simplicity and effectiveness:

Step 1: Set Up Your Tools

First, install the necessary libraries—these will handle text processing, graph building, and visualization:

pip install pandas nltk spacy networkx matplotlib scikit-learn gensim
python -m spacy download en_core_web_sm  # For English text; adjust for other languages

Step 2: Clean & Preprocess Your Text

Raw text has noise that can mess up concept extraction. Let’s standardize it:

import pandas as pd
import nltk
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
import spacy

# Download NLTK resources if you haven't
nltk.download('stopwords')
nltk.download('wordnet')

# Initialize tools
stop_words = set(stopwords.words('english'))
lemmatizer = WordNetLemmatizer()
nlp = spacy.load('en_core_web_sm')

def preprocess_text(text):
    # Lowercase
    text = text.lower()
    # Remove punctuation and extra spaces (using spaCy for better tokenization)
    doc = nlp(text)
    tokens = [token.text for token in doc if not token.is_punct and not token.is_space]
    # Remove stopwords and lemmatize
    cleaned_tokens = [lemmatizer.lemmatize(token) for token in tokens if token not in stop_words]
    return ' '.join(cleaned_tokens)

# Apply preprocessing to your DataFrame's text column (replace 'text' with your column name)
df['cleaned_text'] = df['text'].apply(preprocess_text)

Step 3: Extract Core Concepts

You have two solid options here—pick the one that fits your use case:

Option A: Keyword/Keyphrase Extraction

Extract high-impact terms or noun phrases (these are usually your core concepts):

# Extract noun phrases using spaCy
def extract_concepts(text):
    doc = nlp(text)
    # Filter out short phrases and keep meaningful ones
    concepts = [chunk.text.strip() for chunk in doc.noun_chunks if len(chunk.text.split()) >= 1]
    return concepts

# Add concepts column to DataFrame
df['concepts'] = df['cleaned_text'].apply(extract_concepts)

# Optional: Filter rare concepts to reduce clutter (keep concepts that appear in >= 5% of samples)
concept_counts = pd.Series([item for sublist in df['concepts'] for item in sublist]).value_counts()
min_occurrences = int(len(df) * 0.05)  # Adjust threshold as needed
valid_concepts = set(concept_counts[concept_counts >= min_occurrences].index)
df['filtered_concepts'] = df['concepts'].apply(lambda x: [c for c in x if c in valid_concepts])

Option B: Topic Modeling (LDA)

If you want higher-level thematic concepts, use LDA to extract topics and their top terms:

from sklearn.feature_extraction.text import CountVectorizer
from gensim.models import LdaModel
from gensim.corpora import Dictionary

# Create a dictionary of tokens
texts = [text.split() for text in df['cleaned_text']]
dictionary = Dictionary(texts)
# Filter extreme tokens to reduce noise
dictionary.filter_extremes(no_below=5, no_above=0.8)
corpus = [dictionary.doc2bow(text) for text in texts]

# Train LDA model (adjust num_topics based on your data)
lda_model = LdaModel(corpus=corpus, id2word=dictionary, num_topics=5, random_state=42)

# Extract top terms per topic as core concepts
core_concepts = []
for idx, topic in lda_model.print_topics(-1):
    terms = [term.split('*')[1].strip().strip('"') for term in topic.split('+')]
    core_concepts.extend(terms)
core_concepts = list(set(core_concepts))  # Remove duplicates

Step 4: Build the Undirected Concept Graph

Now, link concepts based on co-occurrence (if they appear in the same text sample) or semantic similarity:

Method 1: Co-Occurrence-Based Edges

import networkx as nx

# Initialize graph
G = nx.Graph()

# Add nodes (filtered concepts from Option A)
for concept in valid_concepts:
    G.add_node(concept, size=concept_counts[concept])

# Add edges based on co-occurrence in the same text
for concepts in df['filtered_concepts']:
    # Create all pairs of concepts in the current text
    for i in range(len(concepts)):
        for j in range(i+1, len(concepts)):
            c1, c2 = concepts[i], concepts[j]
            if G.has_edge(c1, c2):
                # Increase edge weight if already exists
                G[c1][c2]['weight'] += 1
            else:
                G.add_edge(c1, c2, weight=1)

# Optional: Filter weak edges (keep edges with weight >= 2)
edges_to_remove = [(u, v) for u, v, d in G.edges(data=True) if d['weight'] < 2]
G.remove_edges_from(edges_to_remove)

Method 2: Semantic Similarity-Based Edges

If you want edges based on meaning (not just co-occurrence), use word embeddings:

from gensim.models import Word2Vec

# Train Word2Vec on your cleaned text
w2v_model = Word2Vec(texts, vector_size=100, window=5, min_count=5, workers=4)

# Add nodes (core concepts from Option B)
for concept in core_concepts:
    if concept in w2v_model.wv:
        G.add_node(concept)

# Add edges if similarity is above threshold
similarity_threshold = 0.7  # Adjust based on your needs
for i in range(len(core_concepts)):
    for j in range(i+1, len(core_concepts)):
        c1, c2 = core_concepts[i], core_concepts[j]
        if c1 in w2v_model.wv and c2 in w2v_model.wv:
            similarity = w2v_model.wv.similarity(c1, c2)
            if similarity >= similarity_threshold:
                G.add_edge(c1, c2, weight=similarity)

Step 5: Visualize the Graph

Static Visualization (Matplotlib)

import matplotlib.pyplot as plt

# Set node sizes based on occurrence count (from Option A)
node_sizes = [G.nodes[node]['size'] * 20 for node in G.nodes]
# Set edge widths based on weight
edge_widths = [d['weight'] * 0.5 for u, v, d in G.edges(data=True)]

# Draw the graph
plt.figure(figsize=(12, 10))
pos = nx.spring_layout(G, k=0.15)  # Adjust k to spread out nodes
nx.draw_networkx_nodes(G, pos, node_size=node_sizes, node_color='lightblue')
nx.draw_networkx_edges(G, pos, width=edge_widths, edge_color='gray')
nx.draw_networkx_labels(G, pos, font_size=10)
plt.title("Undirected Concept Graph")
plt.axis('off')
plt.show()

Interactive Visualization (Pyvis)

For a zoomable, interactive graph (great for larger datasets):

pip install pyvis
from pyvis.network import Network

net = Network(notebook=True, height="750px", width="100%", bgcolor="#222222", font_color="white")

# Add nodes and edges with properties
for node in G.nodes:
    net.add_node(node, size=G.nodes[node].get('size', 10)*2)
for u, v, d in G.edges(data=True):
    net.add_edge(u, v, width=d['weight'])

# Generate interactive HTML
net.show("concept_graph.html")

Pro Tips

  • Tune Thresholds: Adjust minimum occurrence counts and similarity thresholds to keep the graph readable—too many nodes/edges will make it messy.
  • Customize Concepts: If you have domain-specific terms, add them to a custom stopword list or prioritize them in extraction.
  • Iterate: Start with a small subset of your data to test the pipeline, then scale up to all 100 samples.

内容的提问来源于stack exchange,提问作者Abhishek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:14:38