Microsoft Appsource爬虫无法获取全部数据:如何提取priceModel字段?
解决Power BI Visuals页面priceModel字段提取问题
问题背景
使用Python的BeautifulSoup爬取Microsoft Appsource的Power BI visuals页面时,能正常提取标题、开发者、评分、评分次数,但无法获取标识可视化工具是否免费的priceModel字段,尝试通过div或span元素提取无结果。
解决方案
该页面的产品数据(包括priceModel)是通过动态渲染加载的,实际存储在页面的__NEXT_DATA__脚本标签中,而非直接在静态HTML的可见元素里。因此需要解析该脚本内的JSON数据来获取目标字段:
步骤说明
- 定位页面中id为
__NEXT_DATA__的script标签,该标签包含了所有产品的完整数据。 - 解析标签内的JSON字符串,提取产品列表数据。
- 从每个产品的JSON对象中直接获取priceModel及其他所需字段,无需再逐个查找HTML元素。
修改后的完整代码
import requests from bs4 import BeautifulSoup import json import pandas as pd # 修正base_url,添加页码占位符 base_url = 'https://appsource.microsoft.com/en-us/marketplace/apps?product=power-bi-visuals&page={}' all_data = {'Title': [], 'Owner': [], 'Ratings': [], 'Count of Rates': [], 'Price Model': [], 'Page': []} try: # 示例爬取第1页,可自行调整范围 for page_num in range(1, 2): url = base_url.format(page_num) response = requests.get(url) if response.status_code == 200: # 解析页面,找到存储数据的script标签 soup = BeautifulSoup(response.content, 'html.parser') data_script = soup.find('script', id='__NEXT_DATA__') if data_script: # 解析JSON数据 product_data = json.loads(data_script.text) # 定位产品列表 items = product_data['props']['pageProps']['initialState']['productListing']['items'] for item in items: # 提取所需字段 all_data['Title'].append(item['title']) all_data['Owner'].append(item['publisher']['name']) # 处理评分和评分次数,避免无评分的情况 all_data['Ratings'].append(item.get('rating', '0.0')) all_data['Count of Rates'].append(item.get('ratingCount', 0)) # 提取priceModel字段 all_data['Price Model'].append(item['priceModel']) all_data['Page'].append(page_num) print(f'Page {page_num} content processed') else: print(f'Failed to find data script on page {page_num}') else: print(f'Failed to fetch page {page_num}: {response.status_code}') except Exception as e: print("An error occurred:", str(e)) # 转换为DataFrame并导出 df = pd.DataFrame(all_data) df.to_excel('power_bi_visuals.xlsx', index=False) print('Data written to power_bi_visuals.xlsx')
代码说明
- 修正了原代码中base_url未包含页码占位符的问题,确保分页请求正常工作。
- 通过解析
__NEXT_DATA__脚本中的JSON,直接获取完整产品数据,避免了动态渲染导致的元素无法抓取问题。 - 增加了对无评分产品的兼容处理,避免因字段缺失抛出异常。
- 直接从JSON对象中提取
priceModel字段,该字段值通常为Free或对应的付费模式标识。
内容的提问来源于stack exchange,提问作者OnlySalman
相关产品推荐
相关产品推荐

