如何用Python批量读取关键词并获取Google搜索首条链接?
实现批量关键词Google搜索并保存第一条链接
核心修改思路
把单关键词搜索逻辑封装成可复用的函数,再添加文件读取、批量遍历、结果保存的逻辑,就能实现你要的批量功能。以下是完整实现代码:
完整代码实现
from bs4 import BeautifulSoup import requests import csv import time # 封装单关键词搜索函数,返回第一条链接或None def get_first_google_result(keyword): url = 'https://www.google.com/search' headers = { 'Accept': '*/*', 'Accept-Language': 'en-US,en;q=0.5', 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/98.0.4758.82', } parameters = {'q': keyword} try: # 加延迟避免触发反爬 time.sleep(2) content = requests.get(url, headers=headers, params=parameters).text soup = BeautifulSoup(content, 'html.parser') search_section = soup.find(id='search') if not search_section: return None first_link = search_section.find('a') if first_link and 'href' in first_link.attrs: # 清理Google跳转链接,提取真实URL raw_href = first_link['href'] if raw_href.startswith('/url?q='): return raw_href.split('/url?q=')[1].split('&')[0] return raw_href return None except Exception as e: print(f"搜索关键词 {keyword} 时出错: {str(e)}") return None # 从TXT文件读取关键词(每行一个) def read_keywords_from_txt(file_path): with open(file_path, 'r', encoding='utf-8') as f: return [line.strip() for line in f if line.strip()] # 从CSV文件读取关键词(单列存储,无表头) def read_keywords_from_csv(file_path): keywords = [] with open(file_path, 'r', encoding='utf-8') as f: reader = csv.reader(f) for row in reader: if row: keywords.append(row[0].strip()) return keywords # 保存结果到CSV文件 def save_results_to_csv(results, output_path): with open(output_path, 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(['关键词', '第一条链接']) writer.writerows(results) if __name__ == '__main__': # 选择读取方式:TXT或CSV # keywords = read_keywords_from_txt('keywords.txt') keywords = read_keywords_from_csv('keywords.csv') # 批量处理关键词 results = [] for keyword in keywords: print(f"正在搜索: {keyword}") link = get_first_google_result(keyword) results.append([keyword, link if link else '未找到链接']) # 保存结果 save_results_to_csv(results, 'search_results.csv') print("结果已保存到 search_results.csv")
关键说明
- 反爬处理:添加了
time.sleep(2)延迟,避免短时间内请求过多被Google封禁IP;如果需要更稳定,建议轮换User-Agent或使用代理。 - 链接清理:Google搜索结果的链接是跳转格式(
/url?q=真实链接&...),代码里做了解析提取真实URL。 - 异常处理:捕获请求和解析时的异常,避免单个关键词出错导致程序中断。
- 文件支持:同时支持TXT(每行一个关键词)和CSV(单列无表头)两种关键词输入格式,结果统一保存为CSV方便查看。
使用方法
- 准备关键词文件:
- TXT:每行写一个关键词,保存为
keywords.txt - CSV:单列存储关键词,保存为
keywords.csv
- TXT:每行写一个关键词,保存为
- 代码中选择对应的读取函数,运行即可生成结果文件
search_results.csv
内容的提问来源于stack exchange,提问作者Geovanne Santos
相关产品推荐
相关产品推荐

