如何用Beautiful Soup和Regex从<a>标签href提取应用ID
谷歌应用商店爬虫提取应用ID解决方案
核心思路
Google Play搜索结果里的应用详情页链接格式固定为 /store/apps/details?id=com.facebook.katana,只需从<a>标签的href属性中提取id=后的字符串即可。
具体实现步骤
- 定位目标
<a>标签:优先抓取href包含/store/apps/details?id=的<a>标签(Google的class属性可能动态变化,用href规则匹配更可靠) - 提取应用ID的两种方法:
- 字符串分割法:直接按
id=分割后取后续内容,同时处理可能存在的额外参数href = a_tag.get('href') app_id = href.split('id=')[1].split('&')[0] - 正则匹配法:精准匹配符合应用ID格式的字符串(字母、数字、点号组合)
import re pattern = re.compile(r'id=([a-zA-Z0-9.]+)') match = pattern.search(href) if match: app_id = match.group(1)
- 字符串分割法:直接按
完整示例代码
from bs4 import BeautifulSoup import requests import re # 构造请求(必须加User-Agent避免被拦截) url = "https://play.google.com/store/search?q=facebook&c=apps" headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # 抓取所有应用详情页链接 app_links = soup.find_all('a', href=re.compile(r'/store/apps/details\?id=')) # 批量提取应用ID for link in app_links: href = link.get('href') match = re.search(r'id=([a-zA-Z0-9.]+)', href) if match: print(match.group(1))
注意事项
- 必须添加
User-Agent请求头,否则Google会拦截请求,返回空白页或人机验证页面 - Google Play页面结构可能随时调整,若出现匹配失败,需重新检查页面元素的属性规则
内容的提问来源于stack exchange,提问作者Hemant Kumar Chaudhary
相关产品推荐
相关产品推荐

