如何在Python中使用Mario Vilas谷歌搜索API排除特定网站?
Option 1: Filter Results After Fetching
Since the Mario Vilas googlesearch library doesn’t have native domain exclusion, you can add a simple post-fetch filter to skip unwanted domains. Here’s how to tweak your existing code:
First, define a list of domains you want to block. Then, check each result against this list before adding it to your links collection.
Modified code example:
# List of domains you want to exclude blocked_domains = ["example.com", "unwanted-site.net"] keyword = input("Keyword: ") country = input("Country:") tld_of_country = domain_names[country] links = [] for website in search(keyword, tld=tld_of_country, num=2, stop=2, country="canada", pause=2): # Skip results from blocked domains if not any(domain in website for domain in blocked_domains): try: links.append(website) except: continue
For stricter domain matching (e.g., avoiding partial matches), you can parse the URL to extract the exact domain using urllib.parse—this ensures you don’t accidentally exclude sites that mention the blocked domain in their path.
Option 2: Use an API with Native Domain Exclusion
If you want a more efficient solution without post-processing, consider these alternatives:
- SerpAPI: Offers a dedicated
excludeparameter where you can list domains to block directly in your search request. It handles filtering on the backend, so you get clean results immediately. - Google Custom Search API: Lets you use
siteSearchFilter: "-"alongsidesiteSearchto exclude specific domains. You’ll need to set up an API key and custom search engine, but it’s an official, reliable option.
Both tools let you exclude domains as part of the search query, which is more efficient than filtering results after they’re fetched.
内容的提问来源于stack exchange,提问作者ImNotARobot

