使用Python下载Likee网站视频时生成0字节文件的问题排查求助
Hey there, let's break down why you're hitting that persistent HTTP 204 error and ending up with 0-byte video files. I've dealt with similar media platform scraping issues before, so here's what's likely going on and how to fix it:
1. You're Missing Critical Request Headers
Likee's video delivery servers are probably blocking your requests because they don't look like they're coming from a real browser. When you use requests or urllib directly, you send a minimal request without the headers browsers automatically include—things like User-Agent, Referer, and Accept that tell the server you're a legitimate user.
Fix: Add browser-like headers to your request. Here's an example that mirrors a typical Chrome request:
import requests # Mimic a browser's request headers headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Referer': 'https://likee.video/hashtag/CuteHeadChallenge?lang=en', # Match the page you scraped from 'Accept': 'video/mp4,video/*;q=0.9,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.9' } # Stream the video to handle large files efficiently response = requests.get(data["contentUrl"], headers=headers, stream=True) response.raise_for_status() # Will throw an error if we don't get a 200 OK with open(folderpath + '/vid.mp4', "wb") as file: for chunk in response.iter_content(chunk_size=8192): if chunk: # Skip empty chunks file.write(chunk)
2. Your Requests Aren't Using Valid Session Cookies
Since you used Selenium to load the Likee page, your browser session has cookies that authenticate you as a valid visitor. When you switch to requests, you're starting a fresh, unauthenticated session—so the server rejects your video request with a 204.
Fix: Pass the cookies from your Selenium session to requests:
# Extract cookies from Selenium webdriver selenium_cookies = wd.get_cookies() # Convert to a format requests can use request_cookies = {} for cookie in selenium_cookies: request_cookies[cookie['name']] = cookie['value'] # Now send the request with both headers and cookies response = requests.get( data["contentUrl"], headers=headers, cookies=request_cookies, stream=True ) # ... same download code as above ...
3. The Video URL Has Expired
Looking at your example URL (https://video.like.video/asia_live/2s2/2Dz9d6_4.mp4?crc=1506960817&type=5), the crc parameter is likely a short-lived signature. By the time you switch from extracting the URL with Selenium to downloading it with requests, that signature has expired, leading to the 204 error.
Fix: Combine your extraction and download steps immediately, without unnecessary delays. Avoid long time.sleep() calls after extracting the URL—if you need to wait for elements, use Selenium's WebDriverWait instead of fixed sleeps to minimize lag.
Quick Additional Checks
- If you still get errors, inspect the response headers and content with
print(response.headers)andprint(response.text)—sometimes servers send subtle error messages even with a 204 status. - You can verify your headers/cookies by checking the video request in Chrome DevTools (Network tab, filter for "Media", right-click the video request > Copy > Copy as cURL). Then replicate those headers/cookies in your
requestscode manually.
内容的提问来源于stack exchange,提问作者PerplexedSlime

