Python如何从嵌套数组中批量提取指定位置的车型数据
嵌套数组提取指定字段实现方案
直接通过切片跳过表头、遍历取指定索引值的方式即可完成车型字段提取,以下是爬取场景下的常用实现代码:
Python实现(爬取开发最常用语言)
核心逻辑分两步:
- 对原始嵌套数组做切片
[1:],跳过索引为0的首个['Action']表头数组 - 遍历切片后的所有数据子数组,提取每个子数组索引为3的元素,汇总为结果列表
# 替换为你实际爬取得到的原始嵌套数组即可 raw_data = [['Action'], ['1 796004', '35', '2022-04-28', '2013 FORD FUSION TITANIUM', '43004432', '3FA6P0RU3DR297126', 'CA', 'Copart', 'Batumi, Georgia', 'CAIU7608231EBKG03172414', '2022-05-02', '2022-05-02', '0000-00-00', '', 'dock receipt', 'YES', '', 'No', '', '5/3/2022 Per auction, the title is present and will be prepared for mail out; Follow up for a tracking number-Clara5/9/2022 Per auction, they are still working on mailing out this title; Follow up required for a tracking number-Clara5/11/2022 Per auction, the title was mailed out; tr#776771949089-Clara[Add notes]', 'A779937', '', '', '', '[edit]', ''], ['2 763189', '43', '2022-01-10', '2018 TOYOTA CAMRY', '43241241', '4T1B11HK7JU080162', 'GA', 'Copart', 'Poti, Georgia', 'MRKU5529916217682189', '2022-01-25', '2022-01-28', '2022-06-20', '2022-06-27', 'dock receipt', 'YES', '2022-01-28', 'Yes', '', '[Add notes]', 'A774742', '', '', '', '[edit]', ''], ['3 762850', '37', '2022-01-07', '2017 VOLKSWAGEN TOUAREG', '65835511', 'WVGRF7BP3HD000549', 'CA', 'Copart', 'Batumi, Georgia', 'MSDU7281152EBKG02708589', '2022-02-09', '2022-02-09', '2022-06-07', '2022-06-14', 'dock receipt', 'YES', '2022-01-20', 'Yes', '', '[Add notes]', 'A774650', '', '', '', '[edit]', '']] # 一行代码完成提取 car_model_list = [row[3] for row in raw_data[1:]]
执行后得到的输出结果为:
['2013 FORD FUSION TITANIUM', '2018 TOYOTA CAMRY', '2017 VOLKSWAGEN TOUAREG']
异常兼容写法
如果爬取结果可能存在空行、字段缺失的异常情况,可以增加长度判断避免脚本报错:
car_model_list = [row[3] for row in raw_data[1:] if len(row) >= 4]
JavaScript实现
如果是用Node.js做爬取,逻辑完全一致,对应代码如下:
// 替换为实际爬取的原始数组 const rawData = [['Action'], ['1 796004', '35', '2022-04-28', '2013 FORD FUSION TITANIUM' /* 其余字段省略 */]] // 提取车型列表 const carModelList = rawData.slice(1).map(row => row[3])
内容的提问来源于stack exchange,提问作者Mapper
相关产品推荐
相关产品推荐

