You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas列表列拆分为单行?药房爬虫数据规整与API疑问

一、将DataFrame列表列展开为单行药品信息

你的DataFrame中,除Medicine_Name外其余字段均为列表格式,且每个列表的元素对应同一主药品的不同规格。可以通过以下步骤实现每行对应一条药品规格的格式:

步骤1:批量展开所有列表列

遍历所有列表类型的字段,使用explode方法将列表元素拆分为单独行,同时保持行对应关系:

import pandas as pd
import numpy as np

# 假设你的DataFrame名为df
list_columns = ['Bran_Name_Choices', 'Generic_Name_Choices', 'Manufacture_name', 
                'Manufacture_name_Generic', 'Product_Counry', 'Product_Country_Generic',
                'Prescription_Required', 'Prescription_Required_Generic', 'Product_Details',
                'Product_Details_Generic']

# 批量展开所有列表列
for col in list_columns:
    df = df.explode(col, ignore_index=True)

步骤2:拆分药品详情中的集合数据

Product_Details和Product_Details_Generic中的值是集合,需要拆分为剂量和价格两个独立字段:

# 定义拆分函数:区分剂量(含mg/injection)和价格(含$)
def split_detail(detail_set):
    if pd.isna(detail_set) or not detail_set:
        return pd.Series([np.nan, np.nan])
    dose, price = None, None
    for item in detail_set:
        if 'mg' in item or 'injection' in item:
            dose = item.strip()
        elif '$' in item:
            price = item.strip()
    return pd.Series([dose, price])

# 拆分品牌药和仿制药的详情字段
df[['Brand_Dose', 'Brand_Price']] = df['Product_Details'].apply(split_detail)
df[['Generic_Dose', 'Generic_Price']] = df['Product_Details_Generic'].apply(split_detail)

# 按需删除原集合字段
df = df.drop(columns=['Product_Details', 'Product_Details_Generic'])

步骤3:处理空值(可选)

对仿制药相关的空字段填充默认值,提升数据可读性:

fill_cols = ['Generic_Name_Choices', 'Manufacture_name_Generic', 'Product_Country_Generic',
             'Prescription_Required_Generic', 'Generic_Dose', 'Generic_Price']
df[fill_cols] = df[fill_cols].fillna('N/A')

处理后,每条药品规格会单独占据一行,符合你预期的格式。

二、关于直接调用API爬取数据的可行性

可以通过API直接爬取该网站数据,获取格式更规整的结果,步骤如下:

  1. 抓包分析API:打开浏览器开发者工具(F12)→ 切换到Network标签页,访问搜索页或药品详情页,观察XHR/Fetch类型的请求,寻找返回JSON格式数据的接口(例如搜索接口可能为https://spfpharmacy.com/api/search?drugName=A,详情接口可能为https://spfpharmacy.com/api/drug/[药品ID])。
  2. 模拟请求:复制浏览器请求头(如User-Agent、Cookie),使用requests库直接调用API获取数据:
    import requests
    
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
    }
    
    # 示例:调用搜索API获取字母A的药品列表
    response = requests.get('https://spfpharmacy.com/api/search?drugName=A', headers=headers)
    if response.status_code == 200:
        drug_list = response.json()
        df_api = pd.DataFrame(drug_list)
        print(df_api.head())
    
  3. 注意事项:部分API可能需要携带验证信息(如Cookie),需确保请求头与浏览器一致;同时控制请求频率,避免触发反爬机制。

通过API爬取的原生JSON数据无需处理列表拆分,格式更规整,爬取效率也远高于Selenium。

内容的提问来源于stack exchange,提问作者Lalit Joshi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 23:15:34