基于Python机器学习在Directus中实现自定义过滤搜索的可行性与方案问询
Absolutely feasible! This is a really smart way to combine Directus’s extensibility with machine learning for an intelligent Q&A matching system. Let’s walk through exactly how to build this, step by step:
First, let’s align on the core flow we’ll implement:
- User submits a question to Directus
- Directus saves the question to its database
- Directus triggers a webhook to call your Python ML service
- The ML service searches your pre-built QA library for semantically similar questions
- The service sends the best matching answer back to Directus
- Directus returns the answer to the user (or updates the question entry with the answer for retrieval)
2.1 Set Up Directus for Question Handling
First, get your Directus instance ready to manage questions and trigger the ML workflow:
- Create Collections:
questions: Stores user-submitted questions. Include fields:id(auto-generated primary key),content(text field for the question),created_at(timestamp),matched_answer_id(foreign key linking toqa_pairs), andanswer(text field to store the matched answer).qa_pairs: Your pre-built knowledge base. Include fields:id,question(text),answer(text/rich text), andembedding(JSON field to store ML-generated vector representations of the question).
- Configure Webhook:
Go to Settings → Webhooks → Create Webhook to trigger your Python service when a new question is added:- Name: "Find Similar Answer"
- Trigger: "Create" event on the
questionscollection - URL: The API endpoint of your Python service (e.g.,
http://your-server-ip:5000/find-similar-answer) - Method: POST
- Payload: Select "Send Item Data" to pass the new question’s content and ID to the Python service.
2.2 Build the Python ML Service
This service will handle text embedding generation and similarity search. Here’s how to put it together:
2.2.1 Pick Your Tools
For semantic similarity search (understanding meaning, not just keyword matches), we’ll use:
- Sentence-BERT: A pre-trained model from the
sentence-transformerslibrary to convert text into numerical embeddings (vectors that capture semantic meaning). - Cosine Similarity: To compare the user’s question embedding with embeddings from your QA library and find the closest match.
- Flask: To create a simple API endpoint that Directus can call.
2.2.2 Install Dependencies
Run this in your Python environment:
pip install sentence-transformers flask flask-cors numpy directus-sdk
2.2.3 Write the Service Code
Here’s a working example of the Flask API with ML logic:
from flask import Flask, request, jsonify from flask_cors import CORS from sentence_transformers import SentenceTransformer import numpy as np from sklearn.metrics.pairwise import cosine_similarity from directus_sdk import Directus app = Flask(__name__) CORS(app) # Load pre-trained embedding model (lightweight and effective for most use cases) model = SentenceTransformer('all-MiniLM-L6-v2') # Connect to Directus to fetch QA pairs and update question entries directus = Directus( 'https://your-directus-instance-url.com', token='your-directus-admin-or-service-account-token' ) def get_qa_pairs(): """Fetch all QA pairs and their embeddings from Directus""" response = directus.items('qa_pairs').read(fields=['id', 'question', 'answer', 'embedding']) qa_list = response['data'] # Convert stored JSON embeddings back to numpy arrays embeddings = np.array([item['embedding'] for item in qa_list]) return qa_list, embeddings # Preload QA pairs and embeddings (refresh this if your QA library updates) qa_list, qa_embeddings = get_qa_pairs() @app.route('/find-similar-answer', methods=['POST']) def find_similar_answer(): data = request.get_json() user_question = data.get('content') question_id = data.get('id') if not user_question or not question_id: return jsonify({'error': 'Missing question content or ID'}), 400 # Generate embedding for the user's question user_embedding = model.encode([user_question])[0].reshape(1, -1) # Calculate cosine similarity between user question and all QA pairs similarities = cosine_similarity(user_embedding, qa_embeddings)[0] max_similarity_idx = np.argmax(similarities) best_match = qa_list[max_similarity_idx] similarity_score = float(similarities[max_similarity_idx]) # Set a threshold to avoid returning irrelevant matches (adjust based on testing) if similarity_score < 0.7: answer_text = "Sorry, I couldn't find a matching answer for your question." matched_id = None else: answer_text = best_match['answer'] matched_id = best_match['id'] # Update the question entry in Directus with the result directus.items('questions').update( item_id=question_id, data={ 'matched_answer_id': matched_id, 'answer': answer_text, 'similarity_score': similarity_score } ) return jsonify({'status': 'success', 'answer': answer_text}) if __name__ == '__main__': app.run(host='0.0.0.0', port=5000)
2.2.4 Prepopulate QA Embeddings
Run a one-time script to generate embeddings for your existing QA pairs and save them to Directus:
from sentence_transformers import SentenceTransformer from directus_sdk import Directus model = SentenceTransformer('all-MiniLM-L6-v2') directus = Directus('https://your-directus-instance-url.com', token='your-token') # Fetch all QA pairs without embeddings qa_pairs = directus.items('qa_pairs').read(fields=['id', 'question'])['data'] for pair in qa_pairs: embedding = model.encode(pair['question']).tolist() directus.items('qa_pairs').update(item_id=pair['id'], data={'embedding': embedding}) print("Embeddings generated and saved successfully!")
2.3 Return Results to the User
- If users are submitting questions via Directus’s API, they can either poll the
questionsentry for updates or use Directus’s real-time subscriptions to get the answer instantly once it’s added. - For a smoother experience, build a custom Directus extension endpoint that handles question submission, waits for the ML service to return the answer, and sends it back in a single response.
- Cache Embeddings: If your QA library doesn’t change often, cache embeddings in memory to avoid re-fetching them from Directus on every request.
- Scale with Vector Databases: For large QA libraries (10k+ entries), use tools like FAISS or Pinecone for faster similarity searches instead of in-memory cosine similarity.
- Fine-Tune the Model: If your domain uses specialized terminology, fine-tune the Sentence-BERT model on your QA dataset to improve matching accuracy.
- Adjust Similarity Threshold: Test different threshold values (like 0.6 or 0.8) to balance between missing relevant matches and returning irrelevant ones.
- Add Error Handling: Configure Directus webhook retries for when your Python service is unavailable, and add logging to both systems for debugging.
内容的提问来源于stack exchange,提问作者Cangoo

