Python gRPC流式接口返回13 INTERNAL内部错误求助
问题现象
- 本地打印响应时,内容和类型均显示正常:
Assertion: True Response type: <class 'scrape_pb2.ScrapeResponse'>
- 但在Postman中调用gRPC接口时,仅返回
13 INTERNAL错误,无任何额外提示。 - 无法定位问题根源,也不清楚如何在服务端打印或记录错误信息。
相关代码片段
Proto定义片段
syntax = "proto3"; service ScrapeService { rpc ScrapeSearch(ScrapeRequest) returns (stream ScrapeResponse) {}; } message ScrapeRequest { string url = 1; string keyword = 2; } message ScrapeResponse { oneof result { ScrapeSearchProgress search_progress = 1; ScrapeProductsProgress products_progress = 2; FoundProducts found_products = 3; } } message ScrapeSearchProgress { int32 page = 1; int32 total_products = 2; repeated string product_links = 3; }
scraper.py 核心代码
def get_all_search_products(search_url: str, class_keyword: str): search_driver = webdriver.Firefox(options=options, service=service) search_driver.maximize_window() search_driver.get(search_url) # 抓取第一页 product_links = scrape_search(driver=search_driver, class_keyword=class_keyword) page = 1 search_progress = ScrapeSearchProgress(page=page, total_products=len(product_links), product_links=[]) search_progress.product_links[:] = product_links # 抓取后续页面 while go_to_next_page(search_driver): page += 1 print(f'Scraping page=>{page}') product_links.extend(scrape_search(driver=search_driver, class_keyword=class_keyword)) print(f'Number of products scraped=>{len(product_links)}') search_progress.product_links.extend(product_links) # TODO: 移除该行 if page == 6: break search_progress_response = ScrapeResponse(search_progress=search_progress) yield search_progress_response
服务端代码
class ScrapeService(ScrapeService): def ScrapeSearch(self, request, context): print(f"Request received: {request}") scrape_responses = get_all_search_products(search_url=request.url, class_keyword=request.keyword) for response in scrape_responses: print(f"Assertion: {response.HasField('search_progress')}") print(f"Response type: {type(response)}") yield response
排查与解决建议
1. 给服务端添加错误捕获与日志
通过gRPC的context对象设置错误详情,同时捕获所有异常并打印完整栈信息,这样Postman就能看到具体错误:
import traceback import grpc class ScrapeService(ScrapeService): def ScrapeSearch(self, request, context): print(f"Request received: {request}") try: scrape_responses = get_all_search_products(search_url=request.url, class_keyword=request.keyword) for response in scrape_responses: print(f"Assertion: {response.HasField('search_progress')}") print(f"Response type: {type(response)}") yield response except Exception as e: # 打印完整错误栈到服务端日志 error_detail = f"错误信息: {str(e)}\n{traceback.format_exc()}" print(error_detail) # 将错误详情返回给客户端 context.set_code(grpc.StatusCode.INTERNAL) context.set_details(error_detail) return
2. 修正数据生成逻辑的问题
当前代码存在两个关键逻辑错误:
- 第一页的抓取结果没有
yield返回,客户端会长时间收不到数据触发超时 search_progress.product_links.extend(product_links)会重复追加所有历史链接,导致数据量暴增
修正后的get_all_search_products:
def get_all_search_products(search_url: str, class_keyword: str): search_driver = webdriver.Firefox(options=options, service=service) try: search_driver.maximize_window() search_driver.get(search_url) # 抓取第一页并立即返回 product_links = scrape_search(driver=search_driver, class_keyword=class_keyword) page = 1 search_progress = ScrapeSearchProgress( page=page, total_products=len(product_links), product_links=product_links.copy() ) yield ScrapeResponse(search_progress=search_progress) # 抓取后续页面 while go_to_next_page(search_driver): page += 1 print(f'抓取第{page}页') current_page_links = scrape_search(driver=search_driver, class_keyword=class_keyword) product_links.extend(current_page_links) print(f'累计抓取商品数=>{len(product_links)}') # 返回当前页的新数据(或根据需求返回全部数据) search_progress = ScrapeSearchProgress( page=page, total_products=len(product_links), product_links=current_page_links # 若需返回全部则改为product_links.copy() ) yield ScrapeResponse(search_progress=search_progress) if page == 6: break finally: # 确保浏览器关闭,避免资源泄漏 search_driver.quit()
3. 检查客户端调用配置
- Postman调用gRPC流式接口时,需选择
Streaming Response模式 - 若页面抓取耗时较长,需调整客户端和服务端的超时参数,避免触发超时错误
内容的提问来源于stack exchange,提问作者Yousef
相关产品推荐
相关产品推荐

