You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Microsoft Appsource爬虫无法获取全部数据:如何提取priceModel字段?

解决Power BI Visuals页面priceModel字段提取问题

问题背景

使用Python的BeautifulSoup爬取Microsoft Appsource的Power BI visuals页面时,能正常提取标题、开发者、评分、评分次数,但无法获取标识可视化工具是否免费的priceModel字段,尝试通过div或span元素提取无结果。

解决方案

该页面的产品数据(包括priceModel)是通过动态渲染加载的,实际存储在页面的__NEXT_DATA__脚本标签中,而非直接在静态HTML的可见元素里。因此需要解析该脚本内的JSON数据来获取目标字段:

步骤说明

  1. 定位页面中id为__NEXT_DATA__的script标签,该标签包含了所有产品的完整数据。
  2. 解析标签内的JSON字符串,提取产品列表数据。
  3. 从每个产品的JSON对象中直接获取priceModel及其他所需字段,无需再逐个查找HTML元素。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import json
import pandas as pd

# 修正base_url,添加页码占位符
base_url = 'https://appsource.microsoft.com/en-us/marketplace/apps?product=power-bi-visuals&page={}'

all_data = {'Title': [], 'Owner': [], 'Ratings': [], 'Count of Rates': [], 'Price Model': [], 'Page': []}

try:
    # 示例爬取第1页,可自行调整范围
    for page_num in range(1, 2):
        url = base_url.format(page_num)
        response = requests.get(url)
        
        if response.status_code == 200:
            # 解析页面,找到存储数据的script标签
            soup = BeautifulSoup(response.content, 'html.parser')
            data_script = soup.find('script', id='__NEXT_DATA__')
            
            if data_script:
                # 解析JSON数据
                product_data = json.loads(data_script.text)
                # 定位产品列表
                items = product_data['props']['pageProps']['initialState']['productListing']['items']
                
                for item in items:
                    # 提取所需字段
                    all_data['Title'].append(item['title'])
                    all_data['Owner'].append(item['publisher']['name'])
                    # 处理评分和评分次数,避免无评分的情况
                    all_data['Ratings'].append(item.get('rating', '0.0'))
                    all_data['Count of Rates'].append(item.get('ratingCount', 0))
                    # 提取priceModel字段
                    all_data['Price Model'].append(item['priceModel'])
                    all_data['Page'].append(page_num)
                
                print(f'Page {page_num} content processed')
            else:
                print(f'Failed to find data script on page {page_num}')
        else:
            print(f'Failed to fetch page {page_num}: {response.status_code}')

except Exception as e:
    print("An error occurred:", str(e))

# 转换为DataFrame并导出
df = pd.DataFrame(all_data)
df.to_excel('power_bi_visuals.xlsx', index=False)
print('Data written to power_bi_visuals.xlsx')

代码说明

  • 修正了原代码中base_url未包含页码占位符的问题,确保分页请求正常工作。
  • 通过解析__NEXT_DATA__脚本中的JSON,直接获取完整产品数据,避免了动态渲染导致的元素无法抓取问题。
  • 增加了对无评分产品的兼容处理,避免因字段缺失抛出异常。
  • 直接从JSON对象中提取priceModel字段,该字段值通常为Free或对应的付费模式标识。

内容的提问来源于stack exchange,提问作者OnlySalman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 16:33:16