You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python gRPC流式接口返回13 INTERNAL内部错误求助

问题现象
  • 本地打印响应时,内容和类型均显示正常:
Assertion: True
Response type: <class 'scrape_pb2.ScrapeResponse'>
  • 但在Postman中调用gRPC接口时,仅返回13 INTERNAL错误,无任何额外提示。
  • 无法定位问题根源,也不清楚如何在服务端打印或记录错误信息。

相关代码片段

Proto定义片段

syntax = "proto3";

service ScrapeService {
  rpc ScrapeSearch(ScrapeRequest) returns (stream ScrapeResponse) {};
}

message ScrapeRequest {
  string url = 1;
  string keyword = 2;
}

message ScrapeResponse {
  oneof result {
    ScrapeSearchProgress search_progress = 1;
    ScrapeProductsProgress products_progress = 2;
    FoundProducts found_products = 3;
  }
}

message ScrapeSearchProgress {
  int32 page = 1;
  int32 total_products = 2;
  repeated string product_links = 3;
}

scraper.py 核心代码

def get_all_search_products(search_url: str, class_keyword: str):
    search_driver = webdriver.Firefox(options=options, service=service)
    search_driver.maximize_window()
    search_driver.get(search_url)
    # 抓取第一页
    product_links = scrape_search(driver=search_driver, class_keyword=class_keyword)
    page = 1
    search_progress = ScrapeSearchProgress(page=page, total_products=len(product_links), product_links=[])
    search_progress.product_links[:] = product_links

    # 抓取后续页面
    while go_to_next_page(search_driver):
        page += 1
        print(f'Scraping page=>{page}')
        product_links.extend(scrape_search(driver=search_driver, class_keyword=class_keyword))
        print(f'Number of products scraped=>{len(product_links)}')

        search_progress.product_links.extend(product_links)

        # TODO: 移除该行
        if page == 6:
            break

        search_progress_response = ScrapeResponse(search_progress=search_progress)

        yield search_progress_response

服务端代码

class ScrapeService(ScrapeService):
    def ScrapeSearch(self, request, context):
        print(f"Request received: {request}")
        scrape_responses = get_all_search_products(search_url=request.url, class_keyword=request.keyword)

        for response in scrape_responses:
            print(f"Assertion: {response.HasField('search_progress')}")
            print(f"Response type: {type(response)}")
            yield response

排查与解决建议

1. 给服务端添加错误捕获与日志

通过gRPC的context对象设置错误详情,同时捕获所有异常并打印完整栈信息,这样Postman就能看到具体错误:

import traceback
import grpc

class ScrapeService(ScrapeService):
    def ScrapeSearch(self, request, context):
        print(f"Request received: {request}")
        try:
            scrape_responses = get_all_search_products(search_url=request.url, class_keyword=request.keyword)

            for response in scrape_responses:
                print(f"Assertion: {response.HasField('search_progress')}")
                print(f"Response type: {type(response)}")
                yield response
        except Exception as e:
            # 打印完整错误栈到服务端日志
            error_detail = f"错误信息: {str(e)}\n{traceback.format_exc()}"
            print(error_detail)
            # 将错误详情返回给客户端
            context.set_code(grpc.StatusCode.INTERNAL)
            context.set_details(error_detail)
            return

2. 修正数据生成逻辑的问题

当前代码存在两个关键逻辑错误:

  • 第一页的抓取结果没有yield返回,客户端会长时间收不到数据触发超时
  • search_progress.product_links.extend(product_links)会重复追加所有历史链接,导致数据量暴增

修正后的get_all_search_products:

def get_all_search_products(search_url: str, class_keyword: str):
    search_driver = webdriver.Firefox(options=options, service=service)
    try:
        search_driver.maximize_window()
        search_driver.get(search_url)
        # 抓取第一页并立即返回
        product_links = scrape_search(driver=search_driver, class_keyword=class_keyword)
        page = 1
        search_progress = ScrapeSearchProgress(
            page=page, 
            total_products=len(product_links), 
            product_links=product_links.copy()
        )
        yield ScrapeResponse(search_progress=search_progress)

        # 抓取后续页面
        while go_to_next_page(search_driver):
            page += 1
            print(f'抓取第{page}页')
            current_page_links = scrape_search(driver=search_driver, class_keyword=class_keyword)
            product_links.extend(current_page_links)
            print(f'累计抓取商品数=>{len(product_links)}')

            # 返回当前页的新数据(或根据需求返回全部数据)
            search_progress = ScrapeSearchProgress(
                page=page, 
                total_products=len(product_links), 
                product_links=current_page_links  # 若需返回全部则改为product_links.copy()
            )
            yield ScrapeResponse(search_progress=search_progress)

            if page == 6:
                break
    finally:
        # 确保浏览器关闭,避免资源泄漏
        search_driver.quit()

3. 检查客户端调用配置

  • Postman调用gRPC流式接口时,需选择Streaming Response模式
  • 若页面抓取耗时较长,需调整客户端和服务端的超时参数,避免触发超时错误

内容的提问来源于stack exchange,提问作者Yousef

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 00:05:28