无需Selenium开发Python可执行程序爬取谷歌商家电话号码
解决谷歌商家电话号码爬取问题(无需Selenium)
问题原因
你用lxml无法获取数据,核心原因是谷歌搜索结果的部分内容(包括商家电话)通过JavaScript动态渲染,静态HTTP请求返回的HTML里并没有这些数据,而Selenium能模拟浏览器加载完整页面,所以能拿到结果。
方案一:使用谷歌Places API(推荐)
这是最可靠的官方途径,避免爬虫被封禁,且数据准确。每日50-100次请求完全在免费额度范围内(谷歌免费层每月提供15万次请求)。
步骤:
- 前往谷歌云控制台,申请API密钥,并启用「Places API」服务。
- 编写代码调用API获取商家电话:
import pandas as pd import requests # 替换为你的API密钥 GOOGLE_API_KEY = "你的谷歌Places API密钥" def get_business_phone(business_name, address): # 第一步:搜索商家获取Place ID search_query = f"{business_name} {address}" search_url = f"https://maps.googleapis.com/maps/api/place/textsearch/json?query={search_query}&key={GOOGLE_API_KEY}" search_res = requests.get(search_url).json() if not search_res.get("results"): return None place_id = search_res["results"][0]["place_id"] # 第二步:通过Place ID获取详细信息(包括电话) details_url = f"https://maps.googleapis.com/maps/api/place/details/json?place_id={place_id}&fields=formatted_phone_number&key={GOOGLE_API_KEY}" details_res = requests.get(details_url).json() return details_res.get("result", {}).get("formatted_phone_number") # 示例:从DataFrame批量处理 df = pd.DataFrame({ "商家名称": ["THE ICE PALACE"], "地址": ["1 MAIN WALK CHERRY GROVE NY"] }) df["电话号码"] = df.apply(lambda row: get_business_phone(row["商家名称"], row["地址"]), axis=1) print(df)
方案二:使用JS渲染工具替代Selenium
如果不想用API,可以用requests-html库(基于Pyppeteer,无需手动安装浏览器),它能渲染动态JS内容,且比Selenium更轻量化,适合打包成可执行文件。
代码示例:
from requests_html import HTMLSession import pandas as pd import time def scrape_google_phone(business_name, address): query = f"{business_name} {address} phone" session = HTMLSession() # 设置模拟浏览器的User-Agent,避免被识别为爬虫 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } try: response = session.get(f"https://www.google.com/search?q={query}", headers=headers) # 渲染JS,等待2秒确保页面加载完成 response.html.render(sleep=2) # 尝试用你的XPath提取,或换用更稳定的选择器(比如找带tel:链接的元素) phone_elements = response.html.xpath('//*[@id="rso"]/div[1]/div/div/div/div/div/div[1]/div/div[1]/a/span') if phone_elements: return phone_elements[0].text.strip() # 备选:查找所有带tel:的链接 tel_links = response.html.xpath('//a[contains(@href, "tel:")]') if tel_links: return tel_links[0].text.strip() except Exception as e: print(f"爬取失败:{e}") finally: session.close() # 添加延迟,避免触发谷歌反爬 time.sleep(3) return None # 示例批量处理 df = pd.DataFrame({ "商家名称": ["THE ICE PALACE"], "地址": ["1 MAIN WALK CHERRY GROVE NY"] }) df["电话号码"] = df.apply(lambda row: scrape_google_phone(row["商家名称"], row["地址"]), axis=1) print(df)
打包注意事项:
用PyInstaller打包时,需要包含Pyppeteer的Chromium二进制文件,可通过以下命令:
pyinstaller --onefile --add-binary "{path_to_chromium}";{path_to_chromium} your_script.py
(注:Chromium路径可通过from requests_html import HTMLSession; print(HTMLSession().browser.executable_path)获取)
实用建议
- 无论用哪种方案,都要添加请求延迟(2-3秒/次),避免被谷歌封禁IP。
- 谷歌搜索页面结构经常变化,建议定期检查选择器是否有效,优先用语义化的选择器(比如找
tel:链接)而非固定XPath。 - 若选择API方案,记得监控API使用额度,避免超出免费层产生费用。
内容的提问来源于stack exchange,提问作者Spencer Friedman
相关产品推荐
相关产品推荐

