如何使用Scrapy从字符串中提取BuyPrice与SellPrice?
如何用Scrapy从字符串中提取BuyPrice和SellPrice
嘿,我来帮你搞定这个问题!在Scrapy里提取字符串中的BuyPrice和SellPrice,核心得看你的目标字符串是什么格式——是结构化的JSON、非结构化的纯文本,还是嵌在HTML片段里?下面分三种常见情况给你具体实现方案:
情况1:字符串是结构化格式(比如JSON)
如果你的字符串是标准JSON格式,直接解析成字典是最省心的方式,不用折腾正则或者选择器:
import json from scrapy.spiders import Spider class PriceSpider(Spider): name = 'price_spider' start_urls = ['https://example.com/your-target-page'] def parse(self, response): # 假设从页面响应中拿到目标JSON字符串(比如通过XPath/CSS提取的文本) raw_json_str = '{"BuyPrice": 99.99, "SellPrice": 109.99, "ProductName": "Wireless Headphones"}' # 解析为Python字典 price_dict = json.loads(raw_json_str) # 提取价格字段 buy_price = price_dict.get('BuyPrice') sell_price = price_dict.get('SellPrice') yield { 'BuyPrice': buy_price, 'SellPrice': sell_price }
情况2:字符串是非结构化纯文本
如果是类似「今日行情:BuyPrice: 120.0 USD,SellPrice: 130.5 USD」这种无固定标签的纯文本,用正则提取最直接。Scrapy的Selector支持直接对字符串创建实例,再配合正则方法:
from scrapy.spiders import Spider from scrapy.selector import Selector class PriceSpider(Spider): name = 'price_spider' start_urls = ['https://example.com/your-target-page'] def parse(self, response): # 假设拿到目标纯文本字符串 raw_text = "商品报价:买入价(BuyPrice):50.5元,卖出价(SellPrice):55.0元" # 创建Selector对象处理字符串 sel = Selector(text=raw_text) # 用re_first()提取第一个匹配的价格数值 buy_price = sel.re_first(r'BuyPrice.*?(\d+\.?\d*)') sell_price = sel.re_first(r'SellPrice.*?(\d+\.?\d*)') # 按需转成浮点数类型 if buy_price: buy_price = float(buy_price) if sell_price: sell_price = float(sell_price) yield { 'BuyPrice': buy_price, 'SellPrice': sell_price }
提示:正则可以根据实际文本调整,比如如果价格带货币符号,可改成r'BuyPrice:\s*\$?(\d+\.?\d*)'匹配美元符号的情况
情况3:字符串是HTML片段
如果你的字符串是一段HTML代码(比如从页面中提取的某块带标签的内容),用Scrapy的XPath/CSS选择器会比正则更可靠:
from scrapy.spiders import Spider from scrapy.selector import Selector class PriceSpider(Spider): name = 'price_spider' start_urls = ['https://example.com/your-target-page'] def parse(self, response): # 假设拿到目标HTML片段字符串 html_fragment = '<div class="price-box">BuyPrice: <span class="buy-price">79.99</span> | SellPrice: <span class="sell-price">89.99</span></div>' sel = Selector(text=html_fragment) # 用XPath提取关联文本后的价格 buy_price = sel.xpath('//text()[contains(., "BuyPrice")]/following-sibling::span/text()').get() sell_price = sel.xpath('//text()[contains(., "SellPrice")]/following-sibling::span/text()').get() # 如果有明确的class,也可以用CSS选择器更简洁: # buy_price = sel.css('.buy-price::text').get() # sell_price = sel.css('.sell-price::text').get() yield { 'BuyPrice': float(buy_price) if buy_price else None, 'SellPrice': float(sell_price) if sell_price else None }
另外补充个小技巧:如果目标字符串是直接从页面响应中获取的(比如某个标签的文本),可以直接在response的XPath/CSS中结合正则提取,不用单独创建Selector对象,比如:
buy_price = response.xpath('//div[@class="price-info"]/text()').re_first(r'BuyPrice:\s*(\d+\.?\d*)')
内容的提问来源于stack exchange,提问作者Isaac Grau
相关产品推荐
相关产品推荐

