Python爬取RS电商元器件数据写入CSV时如何去除多余字段名
解决方案
核心逻辑是对抓取到的「字段名+值」拼接文本做拆分,仅提取值部分,推荐使用split(maxsplit=1)做单次分割,避免值本身包含空格导致拆分错误,搭配[-1]索引取最后一段可兼容异常结构场景。
修改后的完整代码:
from bs4 import BeautifulSoup import requests import csv with open('csv/rs_tmp.csv', 'w', newline='') as csvfile: rs_cmp_writer = csv.writer(csvfile) with open('urls_folder/rs_tmp.txt', 'r') as url_file: for line in url_file: url = "https://in.rsdelivers.com/productlist/search?query=" + line.lstrip() cmp_data = [] source = requests.get(url).text soup = BeautifulSoup(source, 'lxml') for snippet in soup.find_all('div', class_='row-inline pill-component-module_pill-component__1WRtv pill-component-module_grey__38ctb'): # 拆分字段名和值,仅保留值,自动兼容带冒号、空格的前缀格式 stock_num_part_num = snippet.text.split(maxsplit=1)[-1].strip(' :') cmp_data.append(stock_num_part_num) for snippet in soup.find_all('div', class_='add-to-basket-cta-component_price-block__1agRR'): # 价格字段同逻辑处理,可按需追加要清理的前缀符号 price = snippet.div.p.text.split(maxsplit=1)[-1].strip(' :₹') cmp_data.append(price) rs_cmp_writer.writerow(cmp_data) print("抓取完成")
说明:
split(maxsplit=1)表示仅按第一个空格拆分字符串为2段,避免值本身带空格被错误拆分[-1]直接取拆分后的最后一段,就算部分页面结构异常没有字段前缀,也不会出现索引报错strip(' :')会自动清除值前后的空格、冒号,适配「RS Stock No.: 12345」这类带冒号的字段格式- 如果有其他需要剔除的固定前缀,直接在
strip的参数里追加对应字符即可
内容的提问来源于stack exchange,提问作者electrophile
相关产品推荐
相关产品推荐

