如何用Pandas将文本批量转换为搜索引擎可查的标题与URL?
批量关键词搜索获取网页标题与URL的Pandas解决方案
一、工具准备
- 需安装的依赖库:
pandas、requests、beautifulsoup4、fake_useragent(用于模拟浏览器请求,降低被拦截概率) - 安装命令:
pip install pandas requests beautifulsoup4 fake_useragent
二、核心流程
- 将关键词列表导入Pandas DataFrame
- 对每个关键词调用搜索引擎,抓取第一条结果的标题和URL
- 把结果存入DataFrame,最终导出为可查看的格式
三、完整代码实现
import pandas as pd import requests from bs4 import BeautifulSoup from fake_useragent import UserAgent import time def fetch_search_result(keyword): # 生成随机浏览器UA,模拟真实请求 ua = UserAgent() headers = {'User-Agent': ua.random} # 构造谷歌搜索URL,可替换为Bing等其他搜索引擎 search_url = f"https://www.google.com/search?q={keyword}" try: # 加入延迟,避免触发反爬机制 time.sleep(1) response = requests.get(search_url, headers=headers) response.raise_for_status() # 解析搜索结果页面 soup = BeautifulSoup(response.text, 'html.parser') first_result = soup.find('div', class_='g') if first_result: # 提取标题和URL title = first_result.find('h3').text if first_result.find('h3') else '无标题' raw_url = first_result.find('a')['href'] if first_result.find('a') else '无URL' # 处理谷歌的跳转URL格式 if raw_url.startswith('/url?q='): clean_url = raw_url.split('/url?q=')[1].split('&')[0] else: clean_url = raw_url return title, clean_url else: return '无匹配结果', '无匹配URL' except Exception as e: print(f"处理关键词 {keyword} 时出错: {str(e)}") return '请求错误', '请求错误' # 读取关键词文件(假设关键词按行存于keywords.txt) keywords_df = pd.read_csv('keywords.txt', header=None, names=['keyword']) # 批量处理所有关键词 keywords_df[['header', 'url']] = keywords_df['keyword'].apply( lambda x: pd.Series(fetch_search_result(x)) ) # 导出结果到CSV文件 keywords_df.to_csv('search_results.csv', index=False) # 打印示例结果 print(keywords_df.to_string(index=False))
四、关键注意事项
- 反爬控制:必须保留
time.sleep(),22000条数据按1秒/条计算约需6小时,可根据实际情况调整延迟时长,避免IP被封禁 - 搜索引擎适配:若无法访问谷歌,可替换为Bing,搜索URL改为
f"https://www.bing.com/search?q={keyword}",同时需调整页面解析的元素类名 - 异常容错:代码内置异常捕获,单个关键词处理失败不会中断整个批量任务
- 合规性:批量搜索需遵守搜索引擎的robots协议,避免过度请求造成服务器负载
内容的提问来源于stack exchange,提问作者Nabih Bawazir
相关产品推荐
相关产品推荐

