Scrapy+Playwright异步集成报错:TypeError协程对象无法转为整数
我正在使用Scrapy结合Playwright加载Google Jobs搜索结果页面,需借助Playwright在浏览器环境中加载页面并点击不同职位查看详情。目标示例URL:https://www.google.com/search?q=product+designer+nyc&ibp=htl;jobs。
在交互式Python环境中代码可正常运行,但Scrapy集成Playwright时遇问题。已正确配置start_requests函数,可通过Playwright打开目标页面。
当前parse函数代码如下:
async def parse(self, response): page = response.meta["playwright_page"] jobs = page.locator("//li") num_jobs = jobs.count() for idx in range(num_jobs): # For each job found, first need to click on it await jobs.nth(idx).click() # Then grab this large section of the page that has details about the job # In that large section, first click a couple of "More" buttons job_details = page.locator("#tl_ditsc") more_button1 = job_details.get_by_text("More job highlights") await more_button1.click() more_button2 = job_details.get_by_text("Show full description") await more_button2.click() # Then take that large section and pass it to another function for parsing soup = BeautifulSoup(job_details, 'html.parser') data = self.parse_single_jd(soup) ... yield {data here} return
运行时在for idx in range(num_jobs)行触发错误:TypeError: 'coroutine' object cannot be interpreted as an integer。推测是对异步parse函数特性理解有误,需让jobs.count()完成求值但无法实现;添加if more_button1.count()检查时也会出现同类错误,寻求解决建议。
异步方法必须加
await调用:Playwright的locator.count()是异步方法,直接调用会返回协程对象而非整数,这是报错的核心原因。修改代码:num_jobs = await jobs.count()检查按钮是否存在时同理:
if await more_button1.count() > 0: await more_button1.click()不能直接将Locator传给BeautifulSoup:
job_details是Playwright的Locator对象,需先获取其HTML内容再解析:job_html = await job_details.inner_html() soup = BeautifulSoup(job_html, 'html.parser')优化循环逻辑(可选):替换索引循环为直接遍历元素句柄,代码更简洁:
async for job_handle in jobs.element_handles(): await job_handle.click() # 后续详情页操作逻辑不变添加动态加载等待:Google Jobs页面元素是动态渲染的,点击职位后需等待详情区域加载完成,避免元素找不到的问题:
# 等待详情区域出现 await page.wait_for_selector("#tl_ditsc", state="visible") job_details = page.locator("#tl_ditsc") # 等待"More"按钮可点击 await more_button1.wait_for(state="enabled") await more_button1.click()
内容的提问来源于stack exchange,提问作者Allen Y

