You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无需Selenium开发Python可执行程序爬取谷歌商家电话号码

解决谷歌商家电话号码爬取问题(无需Selenium)

问题原因

你用lxml无法获取数据,核心原因是谷歌搜索结果的部分内容(包括商家电话)通过JavaScript动态渲染,静态HTTP请求返回的HTML里并没有这些数据,而Selenium能模拟浏览器加载完整页面,所以能拿到结果。

方案一:使用谷歌Places API(推荐)

这是最可靠的官方途径,避免爬虫被封禁,且数据准确。每日50-100次请求完全在免费额度范围内(谷歌免费层每月提供15万次请求)。

步骤:

  1. 前往谷歌云控制台,申请API密钥,并启用「Places API」服务。
  2. 编写代码调用API获取商家电话:
import pandas as pd
import requests

# 替换为你的API密钥
GOOGLE_API_KEY = "你的谷歌Places API密钥"

def get_business_phone(business_name, address):
    # 第一步:搜索商家获取Place ID
    search_query = f"{business_name} {address}"
    search_url = f"https://maps.googleapis.com/maps/api/place/textsearch/json?query={search_query}&key={GOOGLE_API_KEY}"
    search_res = requests.get(search_url).json()
    
    if not search_res.get("results"):
        return None
    
    place_id = search_res["results"][0]["place_id"]
    # 第二步:通过Place ID获取详细信息(包括电话)
    details_url = f"https://maps.googleapis.com/maps/api/place/details/json?place_id={place_id}&fields=formatted_phone_number&key={GOOGLE_API_KEY}"
    details_res = requests.get(details_url).json()
    
    return details_res.get("result", {}).get("formatted_phone_number")

# 示例:从DataFrame批量处理
df = pd.DataFrame({
    "商家名称": ["THE ICE PALACE"],
    "地址": ["1 MAIN WALK CHERRY GROVE NY"]
})

df["电话号码"] = df.apply(lambda row: get_business_phone(row["商家名称"], row["地址"]), axis=1)
print(df)

方案二:使用JS渲染工具替代Selenium

如果不想用API,可以用requests-html库(基于Pyppeteer,无需手动安装浏览器),它能渲染动态JS内容,且比Selenium更轻量化,适合打包成可执行文件。

代码示例:

from requests_html import HTMLSession
import pandas as pd
import time

def scrape_google_phone(business_name, address):
    query = f"{business_name} {address} phone"
    session = HTMLSession()
    # 设置模拟浏览器的User-Agent,避免被识别为爬虫
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    try:
        response = session.get(f"https://www.google.com/search?q={query}", headers=headers)
        # 渲染JS,等待2秒确保页面加载完成
        response.html.render(sleep=2)
        
        # 尝试用你的XPath提取,或换用更稳定的选择器(比如找带tel:链接的元素)
        phone_elements = response.html.xpath('//*[@id="rso"]/div[1]/div/div/div/div/div/div[1]/div/div[1]/a/span')
        if phone_elements:
            return phone_elements[0].text.strip()
        
        # 备选:查找所有带tel:的链接
        tel_links = response.html.xpath('//a[contains(@href, "tel:")]')
        if tel_links:
            return tel_links[0].text.strip()
    except Exception as e:
        print(f"爬取失败:{e}")
    finally:
        session.close()
        # 添加延迟,避免触发谷歌反爬
        time.sleep(3)
    
    return None

# 示例批量处理
df = pd.DataFrame({
    "商家名称": ["THE ICE PALACE"],
    "地址": ["1 MAIN WALK CHERRY GROVE NY"]
})

df["电话号码"] = df.apply(lambda row: scrape_google_phone(row["商家名称"], row["地址"]), axis=1)
print(df)

打包注意事项:

用PyInstaller打包时,需要包含Pyppeteer的Chromium二进制文件,可通过以下命令:

pyinstaller --onefile --add-binary "{path_to_chromium}";{path_to_chromium} your_script.py

(注:Chromium路径可通过from requests_html import HTMLSession; print(HTMLSession().browser.executable_path)获取)

实用建议

  • 无论用哪种方案,都要添加请求延迟(2-3秒/次),避免被谷歌封禁IP。
  • 谷歌搜索页面结构经常变化,建议定期检查选择器是否有效,优先用语义化的选择器(比如找tel:链接)而非固定XPath。
  • 若选择API方案,记得监控API使用额度,避免超出免费层产生费用。

内容的提问来源于stack exchange,提问作者Spencer Friedman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 02:23:15