如何将Pandas列表列拆分为单行?药房爬虫数据规整与API疑问
一、将DataFrame列表列展开为单行药品信息
你的DataFrame中,除Medicine_Name外其余字段均为列表格式,且每个列表的元素对应同一主药品的不同规格。可以通过以下步骤实现每行对应一条药品规格的格式:
步骤1:批量展开所有列表列
遍历所有列表类型的字段,使用explode方法将列表元素拆分为单独行,同时保持行对应关系:
import pandas as pd import numpy as np # 假设你的DataFrame名为df list_columns = ['Bran_Name_Choices', 'Generic_Name_Choices', 'Manufacture_name', 'Manufacture_name_Generic', 'Product_Counry', 'Product_Country_Generic', 'Prescription_Required', 'Prescription_Required_Generic', 'Product_Details', 'Product_Details_Generic'] # 批量展开所有列表列 for col in list_columns: df = df.explode(col, ignore_index=True)
步骤2:拆分药品详情中的集合数据
Product_Details和Product_Details_Generic中的值是集合,需要拆分为剂量和价格两个独立字段:
# 定义拆分函数:区分剂量(含mg/injection)和价格(含$) def split_detail(detail_set): if pd.isna(detail_set) or not detail_set: return pd.Series([np.nan, np.nan]) dose, price = None, None for item in detail_set: if 'mg' in item or 'injection' in item: dose = item.strip() elif '$' in item: price = item.strip() return pd.Series([dose, price]) # 拆分品牌药和仿制药的详情字段 df[['Brand_Dose', 'Brand_Price']] = df['Product_Details'].apply(split_detail) df[['Generic_Dose', 'Generic_Price']] = df['Product_Details_Generic'].apply(split_detail) # 按需删除原集合字段 df = df.drop(columns=['Product_Details', 'Product_Details_Generic'])
步骤3:处理空值(可选)
对仿制药相关的空字段填充默认值,提升数据可读性:
fill_cols = ['Generic_Name_Choices', 'Manufacture_name_Generic', 'Product_Country_Generic', 'Prescription_Required_Generic', 'Generic_Dose', 'Generic_Price'] df[fill_cols] = df[fill_cols].fillna('N/A')
处理后,每条药品规格会单独占据一行,符合你预期的格式。
二、关于直接调用API爬取数据的可行性
可以通过API直接爬取该网站数据,获取格式更规整的结果,步骤如下:
- 抓包分析API:打开浏览器开发者工具(F12)→ 切换到
Network标签页,访问搜索页或药品详情页,观察XHR/Fetch类型的请求,寻找返回JSON格式数据的接口(例如搜索接口可能为https://spfpharmacy.com/api/search?drugName=A,详情接口可能为https://spfpharmacy.com/api/drug/[药品ID])。 - 模拟请求:复制浏览器请求头(如
User-Agent、Cookie),使用requests库直接调用API获取数据:import requests headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 示例:调用搜索API获取字母A的药品列表 response = requests.get('https://spfpharmacy.com/api/search?drugName=A', headers=headers) if response.status_code == 200: drug_list = response.json() df_api = pd.DataFrame(drug_list) print(df_api.head()) - 注意事项:部分API可能需要携带验证信息(如Cookie),需确保请求头与浏览器一致;同时控制请求频率,避免触发反爬机制。
通过API爬取的原生JSON数据无需处理列表拆分,格式更规整,爬取效率也远高于Selenium。
内容的提问来源于stack exchange,提问作者Lalit Joshi
相关产品推荐
相关产品推荐

