You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas将文本批量转换为搜索引擎可查的标题与URL?

批量关键词搜索获取网页标题与URL的Pandas解决方案

一、工具准备

  • 需安装的依赖库:pandas、requests、beautifulsoup4、fake_useragent(用于模拟浏览器请求,降低被拦截概率)
  • 安装命令:
pip install pandas requests beautifulsoup4 fake_useragent

二、核心流程

  1. 将关键词列表导入Pandas DataFrame
  2. 对每个关键词调用搜索引擎,抓取第一条结果的标题和URL
  3. 把结果存入DataFrame,最终导出为可查看的格式

三、完整代码实现

import pandas as pd
import requests
from bs4 import BeautifulSoup
from fake_useragent import UserAgent
import time

def fetch_search_result(keyword):
    # 生成随机浏览器UA,模拟真实请求
    ua = UserAgent()
    headers = {'User-Agent': ua.random}
    
    # 构造谷歌搜索URL,可替换为Bing等其他搜索引擎
    search_url = f"https://www.google.com/search?q={keyword}"
    
    try:
        # 加入延迟,避免触发反爬机制
        time.sleep(1)
        response = requests.get(search_url, headers=headers)
        response.raise_for_status()
        
        # 解析搜索结果页面
        soup = BeautifulSoup(response.text, 'html.parser')
        first_result = soup.find('div', class_='g')
        
        if first_result:
            # 提取标题和URL
            title = first_result.find('h3').text if first_result.find('h3') else '无标题'
            raw_url = first_result.find('a')['href'] if first_result.find('a') else '无URL'
            # 处理谷歌的跳转URL格式
            if raw_url.startswith('/url?q='):
                clean_url = raw_url.split('/url?q=')[1].split('&')[0]
            else:
                clean_url = raw_url
            return title, clean_url
        else:
            return '无匹配结果', '无匹配URL'
    except Exception as e:
        print(f"处理关键词 {keyword} 时出错: {str(e)}")
        return '请求错误', '请求错误'

# 读取关键词文件(假设关键词按行存于keywords.txt)
keywords_df = pd.read_csv('keywords.txt', header=None, names=['keyword'])

# 批量处理所有关键词
keywords_df[['header', 'url']] = keywords_df['keyword'].apply(
    lambda x: pd.Series(fetch_search_result(x))
)

# 导出结果到CSV文件
keywords_df.to_csv('search_results.csv', index=False)

# 打印示例结果
print(keywords_df.to_string(index=False))

四、关键注意事项

  • 反爬控制:必须保留time.sleep(),22000条数据按1秒/条计算约需6小时,可根据实际情况调整延迟时长,避免IP被封禁
  • 搜索引擎适配:若无法访问谷歌,可替换为Bing,搜索URL改为f"https://www.bing.com/search?q={keyword}",同时需调整页面解析的元素类名
  • 异常容错:代码内置异常捕获,单个关键词处理失败不会中断整个批量任务
  • 合规性:批量搜索需遵守搜索引擎的robots协议,避免过度请求造成服务器负载

内容的提问来源于stack exchange,提问作者Nabih Bawazir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 03:42:53