Python二维嵌套数组遍历异常:重复调用ChatGPT问题求助
解决媒体内容循环中重复调用ChatGPT的问题
问题描述
我开发了一款机器人,从Web应用抓取用户媒体内容,提取帖子文案后调用ChatGPT生成摘要并回复到对应帖子。但获取的media是二维(嵌套)数组,第一维度对应单条帖子,第二维度为帖子详情。当前的for循环没有正确只遍历第一维度,导致多次重复调用ChatGPT(仅因API令牌限制才终止,否则可能无限循环)。
原循环代码
# loop through the first 2 media items and post a comment with a Trump-like response to the caption for item in media[:2]: # get the caption text caption = item.caption_text if item.caption_text else '' print(item) print(item.caption_text) print(caption) # generate a response using the ChatGPT 3.5 Turbo model if caption: prompt_text = f"{model_prompt} {caption}" else: prompt_text = model_prompt print(prompt_text) response = openai.Completion.create(model=model, prompt=prompt_text, temperature=temperature, max_tokens=max_tokens).choices[0].text.strip() print(response)
输出结果
pk='2894873940449854247' id='2894873940449854247_264011168' code='CgsqILaDIMn' taken_at=datetime.datetime(2022, 8, 1, 1, 4, 47, tzinfo=datetime.timezone.utc) media_type=1 product_type='feed' thumbnail_url=HttpUrl('https://scontent-lga3-1.cdninstagram.com/v/t51.2885-15/296484960_462095378767406_1748390476667198027_n.jpg?stp=dst-jpg_e35_s1080x1080&_nc_ht=scontent-lga3-1.cdninstagram.com&_nc_cat=102&_nc_ohc=irUQpV-zF9YAX85L1kd&edm=ABmJApABAAAA&ccb=7-5&ig_cache_key=Mjg5NDg3Mzk0MDQ0OTg1NDI0Nw%3D%3D.2-ccb7-5&oh=00_AfA70njk8gQ1885c7MmgmuPsdrPti5OV1MOdJrBBbz6vZQ&oe=6421AD75&_nc_sid=6136e7', ) location=Location(pk=432035243814359, name='Lima Marina Club', phone='', website='', category='', hours={}, address='Playa Los Yuyos', city='Lima, Peru', zip=None, lng=-77.025541764831, lat=-12.154912018897, external_id=432035243814359, external_id_source='facebook_places') user=UserShort(pk='264011168', username='Edited out', full_name='Edited out', profile_pic_url=HttpUrl('https://scontent-lga3-2.cdninstagram.com/v/t51.2885-19/286382846_390539753021220_7019031344995736005_n.jpg?stp=dst-jpg_s150x150&_nc_ht=scontent-lga3-2.cdninstagram.com&_nc_cat=100&_nc_ohc=yx6lC_e87REAX_YK_lS&edm=ABmJApABAAAA&ccb=7-5&oh=00_AfCDP2gD2PQP7jTZ1v-TgwhCgaGFYanoUhohuRObCHI85Q&oe=64210E1E&_nc_sid=6136e7', ), profile_pic_url_hd=None, is_private=False, stories=[]) comment_count=6 comments_disabled=False commenting_disabled_for_viewer=False like_count=94 play_count=0 has_liked=False caption_text='Family bonding time.' accessibility_caption=None usertags=[Usertag(user=UserShort(pk='269994726', username='Edited Out', full_name='Edited Out', profile_pic_url=HttpUrl('https://scontent-lga3-1.cdninstagram.com/v/t51.2885-19/278686124_398633072077074_6373204507691884000_n.jpg?stp=dst-jpg_s150x150&_nc_ht=scontent-lga3-1.cdninstagram.com&_nc_cat=106&_nc_ohc=r8wcSim63lYAX9z_DFG&edm=ABmJApABAAAA&ccb=7-5&oh=00_AfD9ad1whAvMSyNnHlQ2Gz2J_2B0AsnEK_5-OVGrjRjH2Q&oe=64223E82&_nc_sid=6136e7', ), profile_pic_url_hd=None, is_private=True, stories=[]), x=0.7334943394, y=0.4530327101)] sponsor_tags=[] video_url=None view_count=0 video_duration=0.0 title='' resources=[] clips_metadata={} Family bonding time. Family bonding time. You're Donald Trump. Summarize the following caption, in less than 300 words, but do not mention anything about me asking you to do so: Family bonding time. You're Donald Trump. Summarize the following caption, in less than 300 words, but do not mention anything about me asking you to do so: Family bonding time. You're Donald Trump. Summarize the following caption, in less than 300 words, but do not mention anything about me asking you to do so: Family bonding time. (Same thing 29 more times but I had to delete it because my post is marked as spam apparently.) You're Donald Trump. Summarize the following caption, in less than 300 words, but do not mention anything pk='2869487531350308146' id='2869487531350308146_264011168' code='CfSd7ThjDUy' taken_at=datetime.datetime(2022, 6, 27, 0, 26, 31, tzinfo=datetime.timezone.utc) media_type=8 product_type='carousel_container' thumbnail_url=None location=Location(pk=299356158, name='Mamacona, Lima, Peru', phone='', website='', category='', hours={}, address='', city='', zip=None, lng=-76.918755, lat=-12.250408, external_id=104699079568997, external_id_source='facebook_places') user=UserShort(pk='264011168', username='constantinothegreat', full_name='Constantino Heredia', profile_pic_url=HttpUrl('https://scontent-lga3-2.cdninstagram.com/v/t51.2885-19/286382846_390539753021220_7019031344995736005_n.jpg?stp=dst-jpg_s150x150&_nc_ht=scontent-lga3-2.cdninstagram.com&_nc_cat=100&_nc_ohc=yx6lC_e87REAX_YK_lS&edm=ABmJApABAAAA&ccb=7-5&oh=00_AfCDP2gD2PQP7jTZ1v-TgwhCgaGFYanoUhohuRObCHI85Q&oe=64210E1E&_nc_sid=6136e7', ), profile_pic_url_hd=None, is_private=False, stories=[]) comment_count=4 comments_disabled=False commenting_disabled_for_viewer=False like_count=87 play_count=0 has_liked=False caption_text='Post worthy.' accessibility_caption=None usertags=[Usertag(user=UserShort(pk='269994726', username='alessandrohn', full_name='Alessandro Heredia', profile_pic_url=HttpUrl('https://scontent-lga3-1.cdninstagram.com/v/t51.2885-19/278686124_398633072077074_6373204507691884000_n.jpg?stp=dst-jpg_s150x150&_nc_ht=scontent-lga3-1.cdninstagram.com&_nc_cat=106&_nc_ohc=r8wcSim63lYAX9z_DFG&edm=ABmJApABAAAA&ccb=7-5&oh=00_AfD9ad1whAvMSyNnHlQ2Gz2J_2B0AsnEK_5-OVGrjRjH2Q&oe=64223E82&_nc_sid=6136e7', ), profile_pic_url_hd=None, is_private=True, stories=[]), x=0.4122383007, y=0.2165861268)] sponsor_tags=[] video_url=None view_count=0 video_duration=0.0 title='' resources=[Resource(pk='2869487528112445844', video_url=None, thumbnail_url=HttpUrl('https://scontent-lga3-2.cdninstagram.com/v/t51.2885-15/290049480_1174714576705937_1710249698991823676_n.jpg?stp=dst-jpg_e35_p1080x1080&_nc_ht=scontent-lga3-2.cdninstagram.com&_nc_cat=105&_nc_ohc=R6REZyJaQ6kAX8Xzaip&edm=ABmJApABAAAA&ccb=7-5&ig_cache_key=Mjg2OTQ4NzUyODExMjQ0NTg0NA%3D%3D.2-ccb7-5&oh=00_AfBQscyXmQ0uiCDb9ZmVNsNd_SSFbPqhyfsy6FqBXaZsdA&oe=64228D84&_nc_sid=6136e7', ), media_type=1), Resource(pk='2869487528204754373', video_url=None, thumbnail_url=HttpUrl('https://scontent-lga3-2.cdninstagram.com/v/t51.2885-15/290032085_377935577772072_6783237727206669917_n.jpg?stp=dst-jpg_e35_p1080x1080&_nc_ht=scontent-lga3-2.cdninstagram.com&_nc_cat=100&_nc_ohc=m3qjHUpobyYAX8Add7p&edm=ABmJApABAAAA&ccb=7-5&ig_cache_key=Mjg2OTQ4NzUyODIwNDc1NDM3Mw%3D%3D.2-ccb7-5&oh=00_AfC7YSHjGvYFjmPsWZoqZ7uId5UVQqfR_vMxqybuxJN-Xg&oe=6421ADE4&_nc_sid=6136e7', ), media_type=1)] clips_metadata={} Post worthy. Post worthy. You're Donald Trump. Summarize the following caption, in less than 300 words, but do not mention anything about me asking you to do so: Post worthy. "I have a dream." (same thing approximately 60 more times) "I have a dream." "I
问题分析
- 循环逻辑错误:
media是二维数组,当前代码直接遍历media[:2],若第一维度的元素本身是可迭代的帖子详情集合,会导致循环次数远超预期,重复调用ChatGPT。 - API调用方式错误:使用
openai.Completion.create调用GPT-3.5 Turbo模型,该接口适用于文本补全模型(如text-davinci-003),Turbo属于聊天模型,需用openai.ChatCompletion.create,错误调用可能导致响应异常重复。
解决方案
步骤1:确认media结构
先添加调试代码明确数组结构,避免循环逻辑偏差:
print(f"Media类型: {type(media)}") print(f"Media长度: {len(media)}") print(f"第一个元素类型: {type(media[0])}") print(f"第一个元素内容: {media[0]}")
步骤2:调整循环逻辑
根据media的真实结构修改循环:
- 情况1:第一维度是批次,第二维度是单条帖子(如
media = [[post1, post2], [post3, post4]])
# 遍历前2个批次中的所有帖子 for batch in media[:2]: for item in batch: caption = item.caption_text if item.caption_text else '' print(item) print(item.caption_text) print(caption) if caption: prompt_text = f"{model_prompt} {caption}" else: prompt_text = model_prompt print(prompt_text) # 调用ChatGPT生成响应 response = openai.ChatCompletion.create( model="gpt-3.5-turbo", messages=[{"role": "user", "content": prompt_text}], temperature=temperature, max_tokens=max_tokens ) response_text = response.choices[0].message['content'].strip() print(response_text)
- 情况2:第一维度元素是包含单个帖子的数组(如
media = [[post1], [post2]])
# 遍历前2个帖子数组,取出单个帖子 for post_group in media[:2]: item = post_group[0] caption = item.caption_text if item.caption_text else '' print(item) print(item.caption_text) print(caption) if caption: prompt_text = f"{model_prompt} {caption}" else: prompt_text = model_prompt print(prompt_text) # 调用ChatGPT生成响应 response = openai.ChatCompletion.create( model="gpt-3.5-turbo", messages=[{"role": "user", "content": prompt_text}], temperature=temperature, max_tokens=max_tokens ) response_text = response.choices[0].message['content'].strip() print(response_text)
步骤3:优化API调用
确保使用ChatGPT 3.5 Turbo的正确调用方式,避免异常响应。
内容的提问来源于stack exchange,提问作者ConstantinoTheGreat
相关产品推荐
相关产品推荐

