You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于requests与urllib.parse的YouTube搜索爬虫订阅数获取失败问题求助

Fixing YouTube Search Scraper Subscriber Count Retrieval

Hey, I get exactly where you're stuck here—your current code is trying to pull subscriber counts directly from the videoRenderer object in YouTube's search results, but that data just isn't stored there. Subscriber counts are tied to channel entities, not individual video entries, so we need to adjust how we're fetching that info. Let's break this down and fix it step by step.

The Root Problem

YouTube's search result videoRenderer objects only contain basic video metadata (title, URL, channel name, etc.)—they don't include channel-specific stats like subscriber counts. To get that number, you first need to extract the channel ID from the video entry, then match that ID to channel metadata that does include subscriber counts (either from the same search page or by fetching the channel's page directly).

Solution 1: Extract Channel Data from the Same Search Page

YouTube often includes channel cards in the search results (either in the primary or secondary content sections). We can scrape those first to build a map of channel IDs to subscriber counts, then match each video to its channel.

Here's how to modify your _parse_html method:

def _parse_html(self, response):
    results = []
    start = (
        response.index("ytInitialData") + len("ytInitialData") + 3
    )
    end = response.index("};", start) + 1
    json_str = response[start:end]
    data = json.loads(json_str)

    # Step 1: Build a map of channel IDs to subscriber counts from the search page
    channel_subs_map = {}
    primary_contents = data["contents"]["twoColumnSearchResultsRenderer"]["primaryContents"]["sectionListRenderer"]["contents"]

    # Check primary content sections for channel renderers
    for section in primary_contents:
        if "itemSectionRenderer" in section:
            for item in section["itemSectionRenderer"]["contents"]:
                if "channelRenderer" in item:
                    channel_data = item["channelRenderer"]
                    channel_id = channel_data["channelId"]
                    sub_count = channel_data.get("subscriberCountText", {}).get("simpleText", "0 subscribers")
                    channel_subs_map[channel_id] = sub_count

    # Check secondary content (right-side channel recommendations)
    if "secondaryContents" in data["contents"]["twoColumnSearchResultsRenderer"]:
        secondary_contents = data["contents"]["twoColumnSearchResultsRenderer"]["secondaryContents"]["secondarySearchContainerRenderer"]["contents"]
        for item in secondary_contents:
            if "channelRenderer" in item:
                channel_data = item["channelRenderer"]
                channel_id = channel_data["channelId"]
                sub_count = channel_data.get("subscriberCountText", {}).get("simpleText", "0 subscribers")
                channel_subs_map[channel_id] = sub_count

    # Step 2: Process video entries and match to channel data
    videos = primary_contents[0]["itemSectionRenderer"]["contents"]
    for video in videos:
        res = {}
        if "videoRenderer" in video.keys():
            video_data = video.get("videoRenderer", {})
            res["title"] = video_data.get("title", {}).get("runs", [[{}]])[0].get("text", None)
            res["url_suffix"] = video_data.get("navigationEndpoint", {}).get("commandMetadata", {}).get("webCommandMetadata", {}).get("url", None)
            
            # Extract channel ID from the video's byline text
            channel_id = None
            # First check longBylineText (full channel name link)
            byline_runs = video_data.get("longBylineText", {}).get("runs", [])
            for run in byline_runs:
                if "navigationEndpoint" in run and "browseEndpoint" in run["navigationEndpoint"]:
                    channel_id = run["navigationEndpoint"]["browseEndpoint"]["browseId"]
                    break
            # Fallback to ownerText if longBylineText doesn't exist
            if not channel_id:
                owner_runs = video_data.get("ownerText", {}).get("runs", [])
                for run in owner_runs:
                    if "navigationEndpoint" in run and "browseEndpoint" in run["navigationEndpoint"]:
                        channel_id = run["navigationEndpoint"]["browseEndpoint"]["browseId"]
                        break
            
            # Get subscriber count from our map, or mark as unknown
            res["subscribers"] = channel_subs_map.get(channel_id, "Unknown")
            results.append(res)
    return results

Solution 2: Fetch Subscriber Counts Directly from Channel Pages

If a channel doesn't appear as a card in the search results (so our map doesn't have it), we can fetch the channel's page and extract the subscriber count from there. Add this helper method to your class:

def _get_channel_subscribers(self, channel_id):
    BASE_URL = "https://youtube.com"
    channel_url = f"{BASE_URL}/channel/{channel_id}"
    # Mimic a real browser to avoid being blocked
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
    }
    try:
        response = requests.get(channel_url, headers=headers).text
        if "ytInitialData" not in response:
            return "Unknown"
        # Extract the channel's initial data
        start = response.index("ytInitialData") + len("ytInitialData") + 3
        end = response.index("};", start) + 1
        json_str = response[start:end]
        data = json.loads(json_str)
        # Pull subscriber count from the channel header
        sub_count = data["header"]["c4TabbedHeaderRenderer"]["subscriberCountText"]["simpleText"]
        return sub_count
    except (KeyError, json.JSONDecodeError, requests.exceptions.RequestException):
        return "Unknown"

Then update the subscriber count line in _parse_html to use this fallback:

# Get subscriber count from our map, or fetch from channel page
if channel_id in channel_subs_map:
    res["subscribers"] = channel_subs_map[channel_id]
else:
    res["subscribers"] = self._get_channel_subscribers(channel_id)

Important Anti-Scraping Notes

YouTube actively blocks scrapers, so make sure to:

  • Use a realistic User-Agent header (like the one in the helper method)
  • Add delays between requests (use time.sleep(2) or similar) to avoid rate-limiting
  • Avoid scraping at high volumes—YouTube may temporarily block your IP if you're too aggressive

内容的提问来源于stack exchange,提问作者noob at math

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 20:17:39