You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于gRPC与Python抓取Pinterest账户全部聊天记录的技术问题咨询

Fixing Pinterest Full Chat History Retrieval & gRPC Stream Considerations

Let's break down your problem step by step, starting with the immediate issue causing missing chat records, then addressing the gRPC stream question.

1. Root Cause: Broken Chat Collection Logic

The biggest issue right now is your get_chat function—it only returns the first chat message and exits immediately. Here's why:

  • Inside your while True loop, you run return _chat_parser(...) on the first iteration, which exits the function right away.
  • The IndexError catch and msg_counter increment never get a chance to run because the return statement terminates execution.

To fix this, you don't need a clunky while loop—Pinterest already returns all messages in the resource_response.data array. Just parse every item in that array directly:

def get_chat(conversation_id:str, csrftoken:str, _b:str, _pinterest_sess:str) -> List[Dict]:
    options = {"page_size":25,"conversation_id":conversation_id,"no_fetch_context_on_resource":False}
    _cookies = {"csrftoken":csrftoken, "_b":_b, "_pinterest_sess":_pinterest_sess}
    query = {"data": json.dumps({"options":options})}
    encoded_query = urlencode(query).replace("+", "%20")
    url = "{}?{}".format(CHAT_API, encoded_query)
    
    # Fetch and extract all chat data in one go
    response_data = _get(url, _cookies).json()["resource_response"]["data"]
    # Parse every chat message and return as a list
    return [_chat_parser(chat) for chat in response_data]

With this fix, your existing gRPC setup (returning chat_response_array) will start returning all chat records correctly.

2. Should You Use gRPC Streams?

Now, to your question about gRPC stream modes:

  • Do you need streams? If your chat history is very large (hundreds/thousands of messages), streams are better because they avoid sending a huge single payload. For smaller datasets, your existing array-based response works fine.
  • Which stream mode? You want a server-side stream. Here's why:
    • Client sends one request (with auth/conversation ID)
    • Server sends multiple responses (one per chat message) over time
    • This is ideal for returning collections of data where you don't want to wait for all data to be processed before sending results.

Modified Proto for Server-Side Streaming

If you decide to use streams, update your proto definition to replace the array response with a stream:

syntax = "proto3";
service Pinterest{
  // Server-side stream: client sends one request, server streams multiple messages
  rpc GetConversations (request_chat_info) returns (stream chat_response);
}

message request_chat_info{
 string conversation_id = 1;
 string csrftoken = 2;
 string _b = 3;
 string _pinterest_sess = 4;
}

message chat_response{
 string type = 1;
 string id = 2;
 string text = 3;
}

// Remove chat_response_array since we're using streams

Updated gRPC Servicer for Streams

Modify your servicer to yield each chat message instead of returning a single array:

class PinterestService(pb2_grpc.PinterestServicer):
 def GetConversations(self, request, context):
    conversation_id = request.conversation_id
    csrftoken = request.csrftoken
    _b = request._b
    _pinterest_sess = request._pinterest_sess
    chats = _exec(
        get_chat,
        {"conversation_id": conversation_id, "csrftoken": csrftoken, "_b": _b, "_pinterest_sess": _pinterest_sess}
    )
    # Yield each chat message one by one (server-side stream)
    for chat in chats:
        yield pb2.chat_response(**chat)

Final Notes

  • Start by fixing the get_chat function—this will immediately solve your "missing chat records" problem, even without changing to streams.
  • Switch to server-side streams only if you're dealing with large volumes of chat data, as it improves memory usage and latency for both client and server.

内容的提问来源于stack exchange,提问作者Mosihere

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 19:27:42