获取Instagram账号精确粉丝数时遭遇AttributeError错误求助
Instagram粉丝数爬取AttributeError问题解决
问题说明
你尝试通过正则表达式提取Instagram账号粉丝数时触发了AttributeError,错误信息如下:
Exception has occurred: AttributeError
'NoneType' 对象没有属性 'group'
这是因为re.search()没有找到匹配的字符串,返回了None,此时调用.group(1)自然会报错。
核心原因
- Instagram的网页结构可能已更新,你使用的正则表达式不再匹配当前页面中的粉丝数字段
- 直接发送请求未携带浏览器标识,被Instagram的反爬机制拦截,返回的页面内容不包含目标数据
解决方法
1. 先判断匹配结果,避免直接调用group
修改代码,先检查re.search()的返回值是否为None,再进行后续操作:
import requests import re user = "example" url = f'https://www.instagram.com/{user}' # 添加User-Agent模拟浏览器请求 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers).text match_result = re.search(r'"edge_followed_by":{"count":(\d+)}', response) if match_result: follower_count = match_result.group(1) print(f"粉丝数:{follower_count}") else: print("未能找到粉丝数,可能页面结构更新或请求被拦截")
2. 提取共享数据解析(更稳定的方式)
Instagram页面会将用户数据存放在window._sharedData中,提取该数据后用JSON解析,比正则更可靠:
import requests import re import json user = "example" url = f'https://www.instagram.com/{user}' headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers).text # 提取window._sharedData的内容 data_match = re.search(r'window\._sharedData = (.*?);', response, re.DOTALL) if data_match: shared_data = json.loads(data_match.group(1)) # 从JSON结构中提取粉丝数 follower_count = shared_data['entry_data']['ProfilePage'][0]['graphql']['user']['edge_followed_by']['count'] print(f"粉丝数:{follower_count}") else: print("无法获取用户数据,请检查请求头或页面结构")
内容的提问来源于stack exchange,提问作者mcccc
相关产品推荐
相关产品推荐

