You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

urllib.urlopen误将HTTP链接识别为本地文件,YouTube标题爬虫报错

Hey there, let's get this sorted out! The root cause of your error is super clear: the YouTube links in your url-nohttps-list.csv don't have an https:// (or http://) protocol prefix. Python 2.7's urllib.urlopen() treats any string without a protocol header as a local file path, which is why it's trying to open your links like they're files on your Mac instead of web URLs.

Here's a revised version of your script that fixes this issue, plus adds some robustness to handle edge cases:

#!/usr/bin/python
from bs4 import BeautifulSoup
import urllib
import csv

with open('url-titles-list.csv', 'wb') as csv_out:
    fieldnames = ['url', 'title']
    writer = csv.DictWriter(csv_out, fieldnames=fieldnames)
    writer.writeheader()  # Move this outside the loop to avoid duplicate headers
    
    with open('url-nohttps-list.csv', 'rb') as csv_in:
        reader = csv.DictReader(csv_in, fieldnames=['linkurls'])
        
        for row in reader:
            raw_link = row['linkurls'].strip()  # Clean up any extra spaces/newlines
            # Add HTTPS prefix if missing
            if not raw_link.startswith(('http://', 'https://')):
                full_link = 'https://{}'.format(raw_link)
            else:
                full_link = raw_link
            
            try:
                with urllib.urlopen(full_link) as response:
                    html = response.read()
                    soup = BeautifulSoup(html, "html.parser")
                    # Handle cases where the page has no title tag
                    video_title = soup.title.string if soup.title else 'No Title Found'
                    writer.writerow({'url': raw_link, 'title': video_title})
                    print("Successfully fetched title for: {}".format(raw_link))
            except Exception as e:
                error_msg = "Failed to fetch: {}".format(str(e))
                print(error_msg)
                # Log errors to your output CSV instead of crashing
                writer.writerow({'url': raw_link, 'title': error_msg})

Key Fixes & Improvements:

  • Protocol Prefix Addition: The script checks if each link starts with http:// or https://—if not, it adds https:// to turn it into a valid web URL. This stops urllib from treating links as local files.
  • Error Handling: The try-except block prevents a single broken link from crashing the entire scraper. It also logs errors directly into your output CSV so you can review which links failed later.
  • Cleaner Data: strip() removes any accidental spaces or line breaks from your CSV links, which can cause invalid URLs even with the prefix added.
  • Header Fix: Moved writer.writeheader() outside the loop so your output CSV only has one header row instead of repeating it for every entry.

内容的提问来源于stack exchange,提问作者banbourg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:59:43