urllib.urlopen误将HTTP链接识别为本地文件,YouTube标题爬虫报错
Fix: urllib Treating YouTube Links as Local Files in Your Web Scraper
Hey there, let's get this sorted out! The root cause of your error is super clear: the YouTube links in your url-nohttps-list.csv don't have an https:// (or http://) protocol prefix. Python 2.7's urllib.urlopen() treats any string without a protocol header as a local file path, which is why it's trying to open your links like they're files on your Mac instead of web URLs.
Here's a revised version of your script that fixes this issue, plus adds some robustness to handle edge cases:
#!/usr/bin/python from bs4 import BeautifulSoup import urllib import csv with open('url-titles-list.csv', 'wb') as csv_out: fieldnames = ['url', 'title'] writer = csv.DictWriter(csv_out, fieldnames=fieldnames) writer.writeheader() # Move this outside the loop to avoid duplicate headers with open('url-nohttps-list.csv', 'rb') as csv_in: reader = csv.DictReader(csv_in, fieldnames=['linkurls']) for row in reader: raw_link = row['linkurls'].strip() # Clean up any extra spaces/newlines # Add HTTPS prefix if missing if not raw_link.startswith(('http://', 'https://')): full_link = 'https://{}'.format(raw_link) else: full_link = raw_link try: with urllib.urlopen(full_link) as response: html = response.read() soup = BeautifulSoup(html, "html.parser") # Handle cases where the page has no title tag video_title = soup.title.string if soup.title else 'No Title Found' writer.writerow({'url': raw_link, 'title': video_title}) print("Successfully fetched title for: {}".format(raw_link)) except Exception as e: error_msg = "Failed to fetch: {}".format(str(e)) print(error_msg) # Log errors to your output CSV instead of crashing writer.writerow({'url': raw_link, 'title': error_msg})
Key Fixes & Improvements:
- Protocol Prefix Addition: The script checks if each link starts with
http://orhttps://—if not, it addshttps://to turn it into a valid web URL. This stops urllib from treating links as local files. - Error Handling: The
try-exceptblock prevents a single broken link from crashing the entire scraper. It also logs errors directly into your output CSV so you can review which links failed later. - Cleaner Data:
strip()removes any accidental spaces or line breaks from your CSV links, which can cause invalid URLs even with the prefix added. - Header Fix: Moved
writer.writeheader()outside the loop so your output CSV only has one header row instead of repeating it for every entry.
内容的提问来源于stack exchange,提问作者banbourg
相关产品推荐
相关产品推荐

