基于requests与urllib.parse的YouTube搜索爬虫订阅数获取失败问题求助
Hey, I get exactly where you're stuck here—your current code is trying to pull subscriber counts directly from the videoRenderer object in YouTube's search results, but that data just isn't stored there. Subscriber counts are tied to channel entities, not individual video entries, so we need to adjust how we're fetching that info. Let's break this down and fix it step by step.
The Root Problem
YouTube's search result videoRenderer objects only contain basic video metadata (title, URL, channel name, etc.)—they don't include channel-specific stats like subscriber counts. To get that number, you first need to extract the channel ID from the video entry, then match that ID to channel metadata that does include subscriber counts (either from the same search page or by fetching the channel's page directly).
Solution 1: Extract Channel Data from the Same Search Page
YouTube often includes channel cards in the search results (either in the primary or secondary content sections). We can scrape those first to build a map of channel IDs to subscriber counts, then match each video to its channel.
Here's how to modify your _parse_html method:
def _parse_html(self, response): results = [] start = ( response.index("ytInitialData") + len("ytInitialData") + 3 ) end = response.index("};", start) + 1 json_str = response[start:end] data = json.loads(json_str) # Step 1: Build a map of channel IDs to subscriber counts from the search page channel_subs_map = {} primary_contents = data["contents"]["twoColumnSearchResultsRenderer"]["primaryContents"]["sectionListRenderer"]["contents"] # Check primary content sections for channel renderers for section in primary_contents: if "itemSectionRenderer" in section: for item in section["itemSectionRenderer"]["contents"]: if "channelRenderer" in item: channel_data = item["channelRenderer"] channel_id = channel_data["channelId"] sub_count = channel_data.get("subscriberCountText", {}).get("simpleText", "0 subscribers") channel_subs_map[channel_id] = sub_count # Check secondary content (right-side channel recommendations) if "secondaryContents" in data["contents"]["twoColumnSearchResultsRenderer"]: secondary_contents = data["contents"]["twoColumnSearchResultsRenderer"]["secondaryContents"]["secondarySearchContainerRenderer"]["contents"] for item in secondary_contents: if "channelRenderer" in item: channel_data = item["channelRenderer"] channel_id = channel_data["channelId"] sub_count = channel_data.get("subscriberCountText", {}).get("simpleText", "0 subscribers") channel_subs_map[channel_id] = sub_count # Step 2: Process video entries and match to channel data videos = primary_contents[0]["itemSectionRenderer"]["contents"] for video in videos: res = {} if "videoRenderer" in video.keys(): video_data = video.get("videoRenderer", {}) res["title"] = video_data.get("title", {}).get("runs", [[{}]])[0].get("text", None) res["url_suffix"] = video_data.get("navigationEndpoint", {}).get("commandMetadata", {}).get("webCommandMetadata", {}).get("url", None) # Extract channel ID from the video's byline text channel_id = None # First check longBylineText (full channel name link) byline_runs = video_data.get("longBylineText", {}).get("runs", []) for run in byline_runs: if "navigationEndpoint" in run and "browseEndpoint" in run["navigationEndpoint"]: channel_id = run["navigationEndpoint"]["browseEndpoint"]["browseId"] break # Fallback to ownerText if longBylineText doesn't exist if not channel_id: owner_runs = video_data.get("ownerText", {}).get("runs", []) for run in owner_runs: if "navigationEndpoint" in run and "browseEndpoint" in run["navigationEndpoint"]: channel_id = run["navigationEndpoint"]["browseEndpoint"]["browseId"] break # Get subscriber count from our map, or mark as unknown res["subscribers"] = channel_subs_map.get(channel_id, "Unknown") results.append(res) return results
Solution 2: Fetch Subscriber Counts Directly from Channel Pages
If a channel doesn't appear as a card in the search results (so our map doesn't have it), we can fetch the channel's page and extract the subscriber count from there. Add this helper method to your class:
def _get_channel_subscribers(self, channel_id): BASE_URL = "https://youtube.com" channel_url = f"{BASE_URL}/channel/{channel_id}" # Mimic a real browser to avoid being blocked headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } try: response = requests.get(channel_url, headers=headers).text if "ytInitialData" not in response: return "Unknown" # Extract the channel's initial data start = response.index("ytInitialData") + len("ytInitialData") + 3 end = response.index("};", start) + 1 json_str = response[start:end] data = json.loads(json_str) # Pull subscriber count from the channel header sub_count = data["header"]["c4TabbedHeaderRenderer"]["subscriberCountText"]["simpleText"] return sub_count except (KeyError, json.JSONDecodeError, requests.exceptions.RequestException): return "Unknown"
Then update the subscriber count line in _parse_html to use this fallback:
# Get subscriber count from our map, or fetch from channel page if channel_id in channel_subs_map: res["subscribers"] = channel_subs_map[channel_id] else: res["subscribers"] = self._get_channel_subscribers(channel_id)
Important Anti-Scraping Notes
YouTube actively blocks scrapers, so make sure to:
- Use a realistic
User-Agentheader (like the one in the helper method) - Add delays between requests (use
time.sleep(2)or similar) to avoid rate-limiting - Avoid scraping at high volumes—YouTube may temporarily block your IP if you're too aggressive
内容的提问来源于stack exchange,提问作者noob at math

