You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对比本地.txt文章与谷歌获取的网页元描述相似度?

How to Compare Your .txt Article with Google Search Meta Descriptions for Similarity

Let’s walk through the entire workflow with practical, actionable tools and techniques—no overly abstract jargon, just stuff you can implement today:

Step 1: Fetch Google Search Meta Descriptions

First, you need to pull meta descriptions from Google’s search results for your target queries. You have two main options here:

Direct web scraping of Google violates their Terms of Service and will get you blocked quickly. The official API is the safe, reliable way. Here’s how to use it:

  • Set up a project in Google Cloud Console, enable the Custom Search API, and grab an API key + search engine ID.
  • Use Python’s requests library to send API requests. The response will include a snippet field—this is the meta description equivalent in search results.

Example snippet to fetch results:

import requests

API_KEY = "your_api_key"
SEARCH_ENGINE_ID = "your_engine_id"
QUERY = "your target search query"

url = f"https://www.googleapis.com/customsearch/v1?q={QUERY}&key={API_KEY}&cx={SEARCH_ENGINE_ID}"
response = requests.get(url).json()

# Extract meta descriptions (snippets)
meta_descriptions = [item["snippet"] for item in response.get("items", [])]

Option B: Web Scraping (Use with Extreme Caution)

If you can’t use the API, you can scrape Google results, but you’ll need to bypass anti-bot measures:

  • Rotate user agents with the fake_useragent library.
  • Add delays between requests.
  • Use proxies if you’re making a lot of calls.
  • Parse HTML with BeautifulSoup to extract search result snippets or meta[name="description"] tags.

Heads up: This is risky—Google will detect and block you if you’re not careful. Stick to the API whenever possible.

Step 2: Preprocess Your Text

Before calculating similarity, clean both your .txt article and the meta descriptions to ensure fair comparisons:

  • Convert all text to lowercase.
  • Remove punctuation, special characters, and extra whitespace.
  • Strip stopwords (common words like "the" or "and" that don’t add meaningful context) using libraries like nltk or spaCy.
  • Optional: Lemmatize words (reduce to their root form, e.g., "running" → "run") for more accurate matching.

Example preprocessing function with nltk:

import nltk
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
import string

nltk.download('stopwords')
nltk.download('wordnet')

def preprocess_text(text):
    # Lowercase all text
    text = text.lower()
    # Remove punctuation
    text = text.translate(str.maketrans('', '', string.punctuation))
    # Split into individual words
    words = text.split()
    # Remove stopwords and lemmatize
    stop_words = set(stopwords.words('english'))
    lemmatizer = WordNetLemmatizer()
    cleaned_words = [lemmatizer.lemmatize(word) for word in words if word not in stop_words]
    return ' '.join(cleaned_words)

# Process your local article
with open('your_article.txt', 'r', encoding='utf-8') as f:
    article_text = f.read()
cleaned_article = preprocess_text(article_text)

# Process fetched meta descriptions
cleaned_descriptions = [preprocess_text(desc) for desc in meta_descriptions]

Step 3: Calculate Text Similarity

Choose a similarity metric based on your needs—here are the most effective options:

Option 1: TF-IDF + Cosine Similarity (Great for Traditional Word Matching)

TF-IDF converts text into numerical vectors, and cosine similarity measures how aligned those vectors are. This is a lightweight, interpretable solution if you care about word overlap.

Example with sklearn:

from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

# Combine cleaned article and descriptions into one list
all_texts = [cleaned_article] + cleaned_descriptions

# Create TF-IDF vectors
vectorizer = TfidfVectorizer()
tfidf_matrix = vectorizer.fit_transform(all_texts)

# Calculate similarity between the article and each meta description
similarity_scores = cosine_similarity(tfidf_matrix[0:1], tfidf_matrix[1:])[0]

# Print results
for i, score in enumerate(similarity_scores):
    print(f"Meta Description {i+1} Similarity: {score:.4f}")

Option 2: Sentence-BERT (Best for Semantic Similarity)

If you want to capture meaning (not just word overlap), use Sentence-BERT. This pre-trained model generates sentence embeddings that excel at comparing short texts (like meta descriptions) to longer articles.

Example with sentence-transformers:

from sentence_transformers import SentenceTransformer, util

# Load a lightweight pre-trained model
model = SentenceTransformer('all-MiniLM-L6-v2')

# Generate embeddings for the article and descriptions
article_embedding = model.encode(cleaned_article, convert_to_tensor=True)
description_embeddings = model.encode(cleaned_descriptions, convert_to_tensor=True)

# Calculate cosine similarity scores
similarity_scores = util.cos_sim(article_embedding, description_embeddings)[0].tolist()

# Print results
for i, score in enumerate(similarity_scores):
    print(f"Meta Description {i+1} Similarity: {score:.4f}")

Option 3: Jaccard Similarity (Simple for Short Texts)

Jaccard similarity measures the overlap of unique words between two texts. It’s super straightforward and works well for short meta descriptions.

Example:

def jaccard_similarity(text1, text2):
    set1 = set(text1.split())
    set2 = set(text2.split())
    intersection = len(set1 & set2)
    union = len(set1 | set2)
    return intersection / union if union != 0 else 0

# Calculate scores for each meta description
similarity_scores = [jaccard_similarity(cleaned_article, desc) for desc in cleaned_descriptions]

Step 4: Interpret the Results

  • Scores range from 0 (no similarity) to 1 (exact match).
  • Sentence-BERT will give the most accurate semantic similarity scores for most use cases.
  • TF-IDF is better if you need to understand which specific words are driving the similarity.

内容的提问来源于stack exchange,提问作者Mridul Sachan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 02:27:53