Google Finance Stock Screener - Python(Scrapy)爬取股价输出空白问题求解
问题排查及修复方案
核心问题列表
allowed_domains配置错误:该配置仅接受纯域名,不能携带路径,你填写的www.google.com/finance/包含了路径前缀,会导致Scrapy自动过滤所有目标股票页面的请求,没有请求发出自然没有返回结果- 变量名拼写错误:提取市盈率的代码里你写的是
stock_inf[4],变量名少了末尾的o,正确应为stock_info[4],如果请求走到这一步会直接抛出异常 - 默认UA被拦截:Scrapy默认携带的User-Agent会被Google识别为爬虫,直接返回403或者空白内容,没有有效数据可供提取
- 域名匹配问题:你填写的
start_urls和allowed_domains域名不匹配,也会触发请求过滤规则
修复后可运行代码
import scrapy bse_list=['quote/ABB:NSE','quote/AEGISLOG:NSE','quote/AMARAJABAT:NSE','quote/AMBALALSA:NSE','quote/HDFC:NSE','quote/ANDHRAPET:NSE','quote/ANSALAPI:NSE'] class CrawlSpider(scrapy.Spider): name = 'crawl' # 仅填写域名即可,不要带路径 allowed_domains = ['google.com'] # 补全完整的finance域名,避免重定向过滤问题 start_urls = ['https://www.google.com/finance/'] # 配置浏览器UA,避免被拦截,也可以在settings.py里全局设置 custom_settings = { 'USER_AGENT': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } def parse(self, response): for stock in bse_list: url_new = response.urljoin(stock) yield scrapy.Request(url_new, callback = self.parse_book) def parse_book(self, response): stock_name = response.xpath('//*[@class="zzDege"]/text()').extract_first() current_price = response.xpath('//*[@class="YMlKec fxKbKc"]/text()').extract_first() stock_info = response.xpath('//*[@class="P6K39c"]/text()').extract() # 增加非空判断,避免页面结构变动导致索引异常 if len(stock_info) >=5: last_closing_price = stock_info[0] day_range = stock_info[1] year_range = stock_info[2] market_cap = stock_info[3] # 修正变量名拼写错误 p_e_ratio = stock_info[4] yield { "stock_name": stock_name, "current_price": current_price, "last_closing_price": last_closing_price, "day_range": day_range, "year_range": year_range, "market_cap": market_cap, "p_e_ratio": p_e_ratio }
额外注意事项
- 运行时可以加
-s LOG_LEVEL=DEBUG参数查看请求日志,确认请求是否被过滤、返回状态码是否正常 - Google Finance的页面class是动态生成的,后续如果class变更需要重新提取对应元素的定位规则
- 不要短时间发起大量请求,避免触发反爬限制
内容的提问来源于stack exchange,提问作者Anikan
相关产品推荐
相关产品推荐

