使用BeautifulSoup无法获取谷歌完整HTML,如何跨设备爬取赛事日期?
问题:谷歌搜索页面动态内容抓取(足球比赛日期)
我是webscraping新手,想从谷歌搜索页面提取纽卡斯尔下一场足球比赛的日期信息。用BeautifulSoup+requests请求时拿不到完整HTML,推测是谷歌页面用了JavaScript渲染。知道Selenium ChromeDriver能解决,但代码要在其他电脑运行,所以没法用这个方案。
我的代码:
import pandas as pd from bs4 import BeautifulSoup import requests a = "Newcastle" url ="https://www.google.com/search?q=" + a + "+next+match" response = requests.get(url) soup = BeautifulSoup(response.text,"html.parser") print(soup) for a in soup.findAll('div') : print(soup.get_text())
想要定位的元素:
<span class="imso_mh__lr-dt-ds">17/12, 13:30</span>
对应的XPath:
//*[@id="sports-app"]/div/div[3]/div[1]/div/div/div/div/div[1]/div/div[1]/div/span[2]
可行解决方法
用支持JS渲染的无界面HTTP库
推荐requests-html,它基于requests和Pyppeteer,能自动处理JS渲染,首次运行会自动下载轻量Chromium,后续无需额外驱动,打包后可跨电脑运行。示例代码:from requests_html import HTMLSession a = "Newcastle" url = f"https://www.google.com/search?q={a}+next+match" session = HTMLSession() response = session.get(url) # 触发JS渲染 response.html.render() # 定位目标元素 target_span = response.html.find('.imso_mh__lr-dt-ds', first=True) print(target_span.text if target_span else "未找到比赛日期")切换移动端UA请求
谷歌移动端页面的JS渲染逻辑更简单,部分场景下用移动端User-Agent请求,能在静态HTML里拿到目标数据,无需额外依赖:import requests from bs4 import BeautifulSoup a = "Newcastle" url = f"https://www.google.com/search?q={a}+next+match" headers = { 'User-Agent': 'Mozilla/5.0 (iPhone; CPU iPhone OS 16_0 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/16.0 Mobile/15E148 Safari/604.1' } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") target_span = soup.find('span', class_='imso_mh__lr-dt-ds') print(target_span.text if target_span else "未找到比赛日期")直接调用体育数据API
谷歌的比赛数据来自第三方体育数据源,直接用免费公开的足球API获取球队赛程,比依赖谷歌页面结构更稳定,也不用处理动态渲染问题。
内容的提问来源于stack exchange,提问作者flair
相关产品推荐
相关产品推荐

