网页爬取时品牌与机型标题合并,需拆分品牌/机型字段
品牌与机型名称拆分解决方案
问题描述
我正在爬取沃达丰澳大利亚官网的手机品牌、型号、价格等信息,作为爬虫新手完成了基础爬取,但获取的h2标题里品牌和机型名合并成了连续字符串(比如返回的AppleiPhone 14 Pro Max),推测网页HTML未通过父子结构或属性区分二者,希望得到拆分指导。
爬取脚本
import pandas import requests from bs4 import BeautifulSoup url = 'https://www.vodafone.com.au/mobile-phones' html = requests.get(url) soup = BeautifulSoup(html.text, 'html.parser') headers = soup.find_all('h2') divs = soup.find_all('div') model = list(map(lambda h: h.text.strip(), headers)) print(model)
当前爬取返回结果
[ "AppleiPhone 14 Pro Max", "AppleiPhone 14 Pro", "AppleiPhone 14 Plus", "AppleiPhone 14", "SamsungSamsung Galaxy Z Fold4 5G", "SamsungSamsung Galaxy Z Flip4 5G", "SamsungSamsung Galaxy S22 Ultra 5G", "AppleiPhone 13", "GoogleGoogle Pixel 7 Pro", "GoogleGoogle Pixel 7", "AppleiPhone 12", "AppleiPhone 11", "SamsungSamsung Galaxy S22 5G", "SamsungSamsung Galaxy A13 5G", "SamsungSamsung Galaxy A13 4G", "GoogleGoogle Pixel 6 Pro", "GoogleGoogle Pixel 6a", "OPPOOPPO Find X5 Pro 5G", "OPPOOPPO A57 4G", "OPPOOPPO Reno8 5G", "SamsungSamsung Galaxy A53 5G", "SamsungSamsung Galaxy A33 5G", "SamsungSamsung Galaxy Z Fold3 5G", "SamsungSamsung Galaxy S21+ 5G", "SamsungSamsung Galaxy S21 Ultra 5G", "SamsungSamsung Galaxy S21 FE 5G", "AppleiPhone SE (3rd gen)", "TCLTCL 20 Pro 5G", "MotorolaMotorola moto g62 5G", "SamsungSamsung Galaxy A73 5G", "OPPOOPPO Find X5 5G", "OPPOOPPO Find X5 Lite 5G", "MotorolaMotorola moto e22i 4G", "MotorolaMotorola edge 30 pro 5G", "MotorolaMotorola edge 30 5G", "Why choose Vodafone?" ]
理想结果示例
Apple; iPhone 14 Pro Max, Apple; iPhone 14 Pro, Apple; iPhone 14 Plus, Apple; iPhone 14,
拆分方法
针对返回文本的规律,我们可以通过品牌匹配+字符串拆分的方式处理,具体实现如下:
方法思路
观察返回数据可总结出两类模式:
- 品牌名重复两次(如
SamsungSamsung Galaxy...、GoogleGoogle Pixel...) - 品牌名后直接衔接机型名(如
AppleiPhone 14...)
我们可以预先整理品牌列表及其长度,针对不同模式拆分出品牌和机型。
代码实现
在原有脚本基础上添加拆分逻辑:
import pandas import requests from bs4 import BeautifulSoup url = 'https://www.vodafone.com.au/mobile-phones' html = requests.get(url) soup = BeautifulSoup(html.text, 'html.parser') headers = soup.find_all('h2') model = list(map(lambda h: h.text.strip(), headers)) # 定义品牌及对应的长度 brand_info = { "Apple": 5, "Samsung": 7, "Google": 6, "OPPO": 4, "TCL": 3, "Motorola": 8 } def split_brand_model(text): for brand, length in brand_info.items(): # 匹配品牌开头的文本 if text.startswith(brand): # 处理品牌重复的情况(如SamsungSamsung) if text.startswith(brand * 2): return f"{brand}; {text[length:]}" # 处理品牌后直接衔接机型的情况(如AppleiPhone) elif len(text) > length and text[length].isupper(): return f"{brand}; {text[length:]}" # 非手机类标题直接返回原文本 return text # 处理所有机型,过滤掉无关标题 processed_models = [split_brand_model(item) for item in model if not item == "Why choose Vodafone?"] # 输出处理后的结果 for item in processed_models: print(item)
处理后效果
运行上述代码后,将得到类似以下的结果:
Apple; iPhone 14 Pro Max
Apple; iPhone 14 Pro
Apple; iPhone 14 Plus
Apple; iPhone 14
Samsung; Galaxy Z Fold4 5G
...
内容的提问来源于stack exchange,提问作者Karl
相关产品推荐
相关产品推荐

