医生求助:能否直接用现有CSV通过Cypher加载至Neo4j构建药物-酶图谱?
Absolutely! You can load your existing CSV into Neo4j to build your drug-enzyme graph—let’s break this down into simple, actionable steps since you’re new to Cypher.
Step 1: Prep Your CSV Format
First, make sure your CSV has clear, consistent columns defining the source node, target node, and relationship type. Gephi CSVs often follow this structure, but double-check it looks something like this:
source,target,relationship_type DrugA,CYP3A4,INHIBITS DrugB,CYP2D6,ENHANCES CYP3A4,DrugC,METABOLIZES
Quick checks:
- Use all caps for relationship types (like
INHIBITS) to follow Neo4j best practices - Ensure drug and cytochrome names are consistent (no typos like "CYP3a4" vs "CYP3A4")
Step 2: Import CSV into Neo4j
We’ll use Cypher’s LOAD CSV command—it’s straightforward and lets you build nodes and relationships directly in the Neo4j Browser.
First, move your CSV file into Neo4j’s import folder (this is required for the file:/// path to work).
2.1 Create Unique Nodes
Run this query to create all Drug and Cytochrome nodes without duplicates:
// Create Drug and Cytochrome nodes based on relationship direction LOAD CSV WITH HEADERS FROM 'file:///your-drug-enzyme.csv' AS row // Handle drug → cytochrome relationships (INHIBITS/ENHANCES) WHERE row.relationship_type IN ['INHIBITS', 'ENHANCES'] MERGE (drug:Drug {name: row.source}) MERGE (cyto:Cytochrome {name: row.target}) UNION // Handle cytochrome → drug relationships (METABOLIZES) LOAD CSV WITH HEADERS FROM 'file:///your-drug-enzyme.csv' AS row WHERE row.relationship_type = 'METABOLIZES' MERGE (cyto:Cytochrome {name: row.source}) MERGE (drug:Drug {name: row.target})
MERGEensures we only create a node if it doesn’t already exist (prevents duplicates)- Labels (
:Drug,:Cytochrome) categorize your nodes for easier querying later
2.2 Create Relationships
Next, run this query to add the actual connections between nodes:
// Create INHIBITS, ENHANCES, and METABOLIZES relationships LOAD CSV WITH HEADERS FROM 'file:///your-drug-enzyme.csv' AS row // Drug inhibits or enhances cytochrome WHERE row.relationship_type IN ['INHIBITS', 'ENHANCES'] MATCH (drug:Drug {name: row.source}), (cyto:Cytochrome {name: row.target}) MERGE (drug)-[:INHIBITS]->(cyto) WHERE row.relationship_type = 'INHIBITS' MERGE (drug)-[:ENHANCES]->(cyto) WHERE row.relationship_type = 'ENHANCES' UNION // Cytochrome metabolizes drug LOAD CSV WITH HEADERS FROM 'file:///your-drug-enzyme.csv' AS row WHERE row.relationship_type = 'METABOLIZES' MATCH (cyto:Cytochrome {name: row.source}), (drug:Drug {name: row.target}) MERGE (cyto)-[:METABOLIZES]->(drug)
Step 3: Verify Your Graph
After running the queries, check your graph with these simple commands:
- View a sample of nodes:
MATCH (n) RETURN n LIMIT 15 - View a sample of relationships:
MATCH ()-[r]->() RETURN r LIMIT 15 - Run a practical query for your work:
MATCH (c:Cytochrome {name: 'CYP3A4'})<-[:METABOLIZES]-(d:Drug) RETURN d.name(finds all drugs metabolized by CYP3A4)
Pro Tips
- If your CSV has extra attributes (like drug IDs, enzyme gene names), add them to the
MERGEclauses (e.g.,MERGE (drug:Drug {name: row.source, drug_id: row.drug_id})) - For larger datasets, Neo4j Desktop’s Import Tool is faster, but
LOAD CSVis better for learning the basics - Stick with
MERGEinstead ofCREATEeverywhere to avoid duplicate nodes/relationships if you re-run queries later
Once your graph is built, you’ll unlock all the power of a graph database—like exploring complex drug-enzyme interactions that would be hard to spot in a CSV or table.
内容的提问来源于stack exchange,提问作者Robert Alexander

