如何通过ScraperAPI获取完整推文用于情感分析?
问题描述
我正在开展电动汽车情感分析项目,使用ScraperAPI爬取Twitter数据。请求获取10条推文,通过下述代码成功获取,但返回结果中的推文(snippet)被截断。请问是否有方法获取完整推文?是否有其他人遇到过该问题?恳请提供建议,谢谢!
使用的代码
import requests import pandas as pd twitter_data = [] payload = { 'api_key': ' ', # 已填入我的API密钥 'query': 'electric vehicle', 'num': '10' } response = requests.get('https://api.scraperapi.com/structured/twitter/search', params=payload) data = response.json() print(data)
返回的输出(已移除突出显示链接的标题)
{ "search_information": {"query_displayed": "site:twitter.com \"electric vehicle\""}, "organic_results": [ { "position": 0, "title": "Felix Hamer • electricfelix", "snippet": "The Slowdown in US Electric Vehicle Sales Looks More Like a Blip by @tsrandall \"And for all the talk of an EV slowdown, many longer-term...", "highlights": ["Electric Vehicle"] }, { "position": 1, "title": "energy star - X.com", "snippet": "Making the switch to an electric vehicle is not only good for the planet, it could also save you $1400/year on fuel and car maintenance.", "highlights": ["electric vehicle"] }, { "position": 2, "title": "Electric Vehicle Association (@ev_association) / X", "snippet": "We are a network of Electric Vehicle advocates, 100% volunteer led, and represent over 100 chapters and thousands of members across North America. Automotive", "highlights": ["Electric Vehicle"] }, { "position": 3, "title": "US Department of Energy", "snippet": "HOP IN! New electric vehicle charging stations are popping up everywhere, making it easier than ever to power up and ride.", "highlights": ["electric vehicle"] }, { "position": 4, "title": "Big_Orrin", "snippet": "So electric vehicle operating costs continue to use vs ICEs. Price paid per mile the same as ICE cars but ICE cars see their fuel prices...", "highlights": ["electric vehicle"] }, { "position": 5, "title": "Is the Electric Cars Revolution Losing Momentum?", "snippet": "An electric vehicle is probably only suitable for individuals who have a designated area for charging at their homes. Globally, China has...", "highlights": ["electric vehicle"] }, { "position": 6, "title": "Chase Cain", "snippet": "Electric vehicle charging isn't keeping up, frustrating EV drivers ⚡️ according to new report from. @JDPowerAutos.", "highlights": ["Electric vehicle"] }, { "position": 7, "title": "Route Fifty", "snippet": "People who live closer to public electric vehicle chargers view the cars more positively, even when accounting for people's party...", "highlights": ["electric vehicle"] }, { "position": 8, "title": "nyserda", "snippet": "Sunday Read: Electric vehicle fast chargers are growing strongly in the United States with one EV fast charging station for every 15 gas...", "highlights": ["Electric vehicle"] }, { "position": 9, "title": "ICF", "snippet": "The electric vehicle market in the U.S. is growing rapidly, yet EVs are still considered inaccessible to many due to their high upfront cost.", "highlights": ["electric vehicle"] } ], "pagination": { "load_more_url": "http://api.scraperapi.com/?url=https%3A%2F%2Fwww.google.com%2Fsearch%3Fq%3Dsite%3Atwitter.com%2B%2522electric%2Bvehicle%2522%26sca_esv%3Da24d50f4328fb081%26gl%3DUS%26ei%3DUDBXZpXaONK9kPIP_6eTSA%26start%3D10%26sa%3DN&autoparse=true", "pages_count": 1 } }
解决方案与建议
关于snippet被截断的原因
你当前调用的是ScraperAPI的结构化Twitter搜索接口,返回的snippet本质是Google搜索结果的摘要,而非Twitter原生的完整推文内容,所以会被Google自动截断。不少用户都遇到过这个问题——因为该接口是基于Google索引的Twitter内容做的结构化提取,而非直接对接Twitter API获取原生数据。
获取完整推文的可行方法
- 切换到ScraperAPI的Twitter原生数据接口
检查ScraperAPI官方文档,确认是否有直接获取单条推文详情的接口。如果有,可以通过当前结果中的用户信息拼接推文URL,或者接口返回的链接字段,调用详情接口拉取完整内容。 - 直接抓取推文原生页面
从搜索结果中提取每条推文的实际URL(比如从用户名和推文ID拼接),使用ScraperAPI的通用网页抓取功能请求推文原生页面,再用BeautifulSoup等工具从HTML中解析完整推文内容。示例代码如下:# 假设已从organic_results中拿到推文url tweet_url = "https://x.com/electricfelix/status/xxxxxx" payload = { 'api_key': '你的API密钥', 'url': tweet_url } response = requests.get('https://api.scraperapi.com/', params=payload) # 后续用解析工具提取完整推文内容 - 直接使用Twitter官方API
如果项目允许,申请Twitter(X)的官方API(比如X API v2),可直接获取完整推文内容,还能附带点赞数、转发数等元数据,更适配情感分析场景。缺点是需要申请权限,有调用额度限制。
额外建议
- 先确认ScraperAPI的结构化Twitter搜索接口是否支持返回完整推文的字段,若支持,只需在请求参数中添加对应标识(如
include_full_text=true,具体以文档为准)。 - 若采用网页抓取方式,注意Twitter页面结构可能变动,需定期调整解析逻辑,ScraperAPI会帮你处理大部分反爬机制。
内容的提问来源于stack exchange,提问作者Glare07
相关产品推荐
相关产品推荐

