You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup和Regex从<a>标签href提取应用ID

谷歌应用商店爬虫提取应用ID解决方案

核心思路

Google Play搜索结果里的应用详情页链接格式固定为 /store/apps/details?id=com.facebook.katana,只需从<a>标签的href属性中提取id=后的字符串即可。

具体实现步骤

  • 定位目标<a>标签:优先抓取href包含/store/apps/details?id=的<a>标签(Google的class属性可能动态变化,用href规则匹配更可靠)
  • 提取应用ID的两种方法:
    1. 字符串分割法:直接按id=分割后取后续内容,同时处理可能存在的额外参数
      href = a_tag.get('href')
      app_id = href.split('id=')[1].split('&')[0]
      
    2. 正则匹配法:精准匹配符合应用ID格式的字符串(字母、数字、点号组合)
      import re
      pattern = re.compile(r'id=([a-zA-Z0-9.]+)')
      match = pattern.search(href)
      if match:
          app_id = match.group(1)
      

完整示例代码

from bs4 import BeautifulSoup
import requests
import re

# 构造请求(必须加User-Agent避免被拦截)
url = "https://play.google.com/store/search?q=facebook&c=apps"
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

# 抓取所有应用详情页链接
app_links = soup.find_all('a', href=re.compile(r'/store/apps/details\?id='))

# 批量提取应用ID
for link in app_links:
    href = link.get('href')
    match = re.search(r'id=([a-zA-Z0-9.]+)', href)
    if match:
        print(match.group(1))

注意事项

  • 必须添加User-Agent请求头,否则Google会拦截请求,返回空白页或人机验证页面
  • Google Play页面结构可能随时调整,若出现匹配失败,需重新检查页面元素的属性规则

内容的提问来源于stack exchange,提问作者Hemant Kumar Chaudhary

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 05:20:29