基于gRPC与Python抓取Pinterest账户全部聊天记录的技术问题咨询
Let's break down your problem step by step, starting with the immediate issue causing missing chat records, then addressing the gRPC stream question.
1. Root Cause: Broken Chat Collection Logic
The biggest issue right now is your get_chat function—it only returns the first chat message and exits immediately. Here's why:
- Inside your
while Trueloop, you runreturn _chat_parser(...)on the first iteration, which exits the function right away. - The
IndexErrorcatch andmsg_counterincrement never get a chance to run because the return statement terminates execution.
To fix this, you don't need a clunky while loop—Pinterest already returns all messages in the resource_response.data array. Just parse every item in that array directly:
def get_chat(conversation_id:str, csrftoken:str, _b:str, _pinterest_sess:str) -> List[Dict]: options = {"page_size":25,"conversation_id":conversation_id,"no_fetch_context_on_resource":False} _cookies = {"csrftoken":csrftoken, "_b":_b, "_pinterest_sess":_pinterest_sess} query = {"data": json.dumps({"options":options})} encoded_query = urlencode(query).replace("+", "%20") url = "{}?{}".format(CHAT_API, encoded_query) # Fetch and extract all chat data in one go response_data = _get(url, _cookies).json()["resource_response"]["data"] # Parse every chat message and return as a list return [_chat_parser(chat) for chat in response_data]
With this fix, your existing gRPC setup (returning chat_response_array) will start returning all chat records correctly.
2. Should You Use gRPC Streams?
Now, to your question about gRPC stream modes:
- Do you need streams? If your chat history is very large (hundreds/thousands of messages), streams are better because they avoid sending a huge single payload. For smaller datasets, your existing array-based response works fine.
- Which stream mode? You want a server-side stream. Here's why:
- Client sends one request (with auth/conversation ID)
- Server sends multiple responses (one per chat message) over time
- This is ideal for returning collections of data where you don't want to wait for all data to be processed before sending results.
Modified Proto for Server-Side Streaming
If you decide to use streams, update your proto definition to replace the array response with a stream:
syntax = "proto3"; service Pinterest{ // Server-side stream: client sends one request, server streams multiple messages rpc GetConversations (request_chat_info) returns (stream chat_response); } message request_chat_info{ string conversation_id = 1; string csrftoken = 2; string _b = 3; string _pinterest_sess = 4; } message chat_response{ string type = 1; string id = 2; string text = 3; } // Remove chat_response_array since we're using streams
Updated gRPC Servicer for Streams
Modify your servicer to yield each chat message instead of returning a single array:
class PinterestService(pb2_grpc.PinterestServicer): def GetConversations(self, request, context): conversation_id = request.conversation_id csrftoken = request.csrftoken _b = request._b _pinterest_sess = request._pinterest_sess chats = _exec( get_chat, {"conversation_id": conversation_id, "csrftoken": csrftoken, "_b": _b, "_pinterest_sess": _pinterest_sess} ) # Yield each chat message one by one (server-side stream) for chat in chats: yield pb2.chat_response(**chat)
Final Notes
- Start by fixing the
get_chatfunction—this will immediately solve your "missing chat records" problem, even without changing to streams. - Switch to server-side streams only if you're dealing with large volumes of chat data, as it improves memory usage and latency for both client and server.
内容的提问来源于stack exchange,提问作者Mosihere

