在R语言中基于HR与邮件数据区分节点内部/外部联系的技术问询
Hey there! Let's work through this step by step to get your graph set up right and unlock useful insights from your HR and email data.
First, let's make sure your graph g1 includes the department (and other HR) attributes from df1. If you didn't include the vertices parameter when creating the graph initially, you can either re-create the graph properly or add the attributes afterward:
Option 1: Re-create the graph with HR data included
# Assuming df2 has "sender" and "receiver" columns; df1 has "ID" (email) and "department" columns g1 <- graph.data.frame(df2[, c("sender", "receiver")], directed = TRUE, vertices = df1)
The vertices argument automatically maps df1's rows to the graph nodes using the matching email addresses (since df1$ID matches the sender/receiver values in df2). Now every node in g1 will have a department attribute (plus any other fields from df1).
Option 2: Add department attributes to an existing graph
If you already have g1 created without the HR data, run this to attach the department info:
# Get all node email addresses from the graph node_emails <- V(g1)$name # Match each node's email to df1's ID and pull the corresponding department V(g1)$department <- df1$department[match(node_emails, df1$ID)]
Since you mentioned multiple emails between the same sender-receiver pairs, it's useful to turn these into weighted edges where the weight is the number of emails exchanged:
# First, aggregate the email data to count emails per sender-receiver pair library(dplyr) edge_counts <- df2 %>% group_by(sender, receiver) %>% summarise(email_count = n(), .groups = "drop") # Create a weighted graph with these counts (and still include HR data) g1_weighted <- graph.data.frame(edge_counts, directed = TRUE, vertices = df1) # Verify the edge weights are set correctly E(g1_weighted)$email_count
Now that your graph is properly set up, here are a few common tasks you might want to run:
- Calculate email activity metrics:
# Unweighted out-degree (number of unique people a node sent emails to) V(g1)$unique_recipients <- degree(g1, mode = "out") # Weighted out-strength (total number of emails sent) V(g1_weighted)$total_emails_sent <- strength(g1_weighted, mode = "out", weights = E(g1_weighted)$email_count) - Analyze cross-department communication:
# Extract edges with sender/receiver departments edge_departments <- data.frame( sender_email = E(g1_weighted)$sender, sender_dept = V(g1_weighted)$department[match(E(g1_weighted)$sender, V(g1_weighted)$name)], receiver_email = E(g1_weighted)$receiver, receiver_dept = V(g1_weighted)$department[match(E(g1_weighted)$receiver, V(g1_weighted)$name)], email_count = E(g1_weighted)$email_count ) # Count total emails between department pairs cross_dept_comm <- edge_departments %>% group_by(sender_dept, receiver_dept) %>% summarise(total_emails = sum(email_count), .groups = "drop")
内容的提问来源于stack exchange,提问作者shima yaghoubian

