Scrapy爬取Audible无输出求助:生成的CSV文件为空
解决Scrapy爬取Audible搜索页生成空CSV的问题
可能原因及对应解决方案
1. 反爬拦截:请求缺少必要标识或被识别为爬虫
Audible的反爬机制仅靠User-Agent无法绕过,需要补充浏览器常用请求头,或启用动态页面渲染来模拟真实用户行为。
补充完整请求头
修改start_requests和分页请求的Headers,添加更多浏览器原生字段:
def start_requests(self): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.audible.com/', 'Connection': 'keep-alive', 'Upgrade-Insecure-Requests': '1' } yield scrapy.Request( url='https://www.audible.com/search/', callback=self.parse, headers=headers ) # 下一页请求复用同一份Headers if next_page_url: yield response.follow( url=next_page_url, callback=self.parse, headers=headers )
启用Playwright渲染动态内容
如果页面内容由JavaScript动态加载,Scrapy静态请求无法获取数据。先安装依赖:
pip install scrapy-playwright
在项目settings.py中配置:
DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, "args": ["--no-sandbox", "--disable-dev-shm-usage"], }
修改请求启用Playwright:
def start_requests(self): yield scrapy.Request( url='https://www.audible.com/search/', callback=self.parse, meta={"playwright": True}, headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36'} ) # 下一页同样启用Playwright if next_page_url: yield response.follow( url=next_page_url, callback=self.parse, meta={"playwright": True}, headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/115.0.0.0 Safari/537.36'} )
2. XPath选择器失效
Audible页面结构可能已更新,原选择器无法定位元素。通过浏览器开发者工具重新定位:
修正产品容器及字段XPath
def parse(self, response): # 替换为更稳定的产品列表项选择器 product_container = response.xpath('//li[contains(@class, "productListItem")]') for product in product_container: # 修正标题选择器 book_title = product.xpath('.//h3[@class="bc-heading bc-color-base"]/a/text()').get() # 修正作者选择器 book_author = product.xpath('.//li[contains(@class, "authorLabel")]/span[2]/a/text()').getall() # 修正时长选择器 book_length = product.xpath('.//li[contains(@class, "runtimeLabel")]/span[2]/text()').get() book_author_string = ' '.join(book_author) yield { 'title': book_title, 'author': book_author_string, 'length': book_length, 'User-Agent': response.request.headers['User-Agent'] } self.logger.debug("Scraped item: %s", {'title': book_title, 'author': book_author_string, 'length': book_length}) # 修正分页选择器 next_page_url = response.xpath('//span[@class="nextButton"]/a/@href').get() if next_page_url: yield response.follow(url=next_page_url, callback=self.parse, headers=headers)
3. 检查Robots.txt限制
若Audible的robots.txt禁止爬取/search/路径,需在settings.py中关闭遵守规则(仅用于教育学习):
ROBOTSTXT_OBEY = False
4. 调试请求响应
在parse方法开头添加调试代码,确认请求是否成功返回有效内容:
def parse(self, response): self.logger.info(f"Response status code: {response.status}") # 保存响应到本地,查看实际页面内容 with open('audible_response.html', 'wb') as f: f.write(response.body) # 后续解析代码...
验证步骤
- 开启调试日志运行爬虫,查看是否有爬取条目输出:
scrapy crawl audible --logfile=scrapy.log -s LOG_LEVEL=DEBUG
- 检查生成的
audible_response.html,确认页面是否正常加载并包含产品数据。
内容的提问来源于stack exchange,提问作者Sazzad Hossain Nibir
相关产品推荐
相关产品推荐

