如何通过爬取网站筛选相似句子?现有doc2vec方案是否有更优替代?
Great question! Your current plan of scraping job site FAQs and using doc2vec + clustering is a solid baseline, but modern NLP tools offer more accurate and flexible alternatives tailored to semantic similarity tasks. Here are the top optimized paths to consider:
1. Use Pre-trained Sentence Embedding Models (Better Than Doc2Vec)
Doc2vec works for basic cases, but transformer-based sentence embedding models like Sentence-BERT, E5, or MiniLM are far better at capturing nuanced context—critical for matching questions like "How much time will interview take?" with "Duration of the interview."
- These models come pre-trained on massive text corpora, so you don’t need to build a model from scratch. For your hiring FAQ use case, you can even fine-tune them on your scraped data to boost domain-specific accuracy.
- Instead of clustering, pair these embeddings with open-source vector databases (like FAISS or Annoy) to run real-time approximate nearest neighbor searches. This lets you dynamically add new FAQs without re-running full clustering, and get instant similar sentence results for any input query.
2. Augment Your Dataset with LLM-Generated Synonyms
If scraping 30-40 job sites feels labor-intensive, or your dataset is small, use a lightweight LLM to generate additional similar questions for your existing FAQs. For example:
Input: "How much time will interview take?"
Generated variations: "What's the expected length of the interview?" "How long should I set aside for the interview?"
This expands your dataset, making your embedding models more robust to the different phrasing styles common in hiring-related questions.
3. Hybrid Rule + Semantic Matching for Efficiency
For larger datasets, combine rule-based filtering with semantic embedding to speed up queries:
- First, use simple rules (e.g., keyword matching for "interview" + "duration/length/time") to narrow down potential candidates.
- Then run semantic similarity checks on the filtered subset using your sentence embedding model. This cuts down on computation time while maintaining accuracy.
4. Domain-Specific Fine-Tuning
If you have enough scraped FAQ data (10k+ entries), fine-tuning a pre-trained sentence embedding model on your hiring-specific QA pairs will yield the most precise results. This teaches the model to prioritize the exact phrasing and context unique to job interview questions, rather than relying on general text patterns.
How This Compares to Your Original Plan
Your doc2vec + clustering approach is static—adding new FAQs requires re-clustering, and it struggles with subtle semantic differences. The alternatives above are:
- More accurate: Better at understanding context and synonymy.
- More flexible: Support dynamic updates and real-time queries.
- Less labor-intensive: Pre-trained models reduce the need for massive training data.
Pick the approach that aligns with your available data and compute resources: start with Sentence-BERT + FAISS if you want a quick, high-quality solution, or invest in fine-tuning if you need domain-specific precision.
内容的提问来源于stack exchange,提问作者Krishna Mohan Varma Dalapati

