Python3下载PDF时报TypeError: expected string or bytes-like object错误求助
解决Python下载PDF时的TypeError问题
Hey there, let's break down what's going wrong here and fix it step by step!
直接错误原因
The TypeError you're seeing comes down to one critical mistake in your getmth_lk method:
You're passing rpLk (an empty list you initialized) to urllib.request.urlretrieve() instead of the actual PDF URL (url_) from your loop. The urlretrieve() function expects a string/bytes URL, not a list—hence the error.
其他需要修复的问题
Beyond that, there are a few other issues in your code that will cause errors or unexpected behavior:
- Missing imports: You're using
requests,BeautifulSoup,urllib, andpathlibbut haven't imported them. - Undefined variables:
reportLinksandroot_urlaren't defined anywhere in your code. - Incorrect list appending: You're trying to append lists directly to
rpLkinstead of extending it (or using it correctly). - Path handling: Hardcoding
/home/might not be portable, and you're not using the path object you created properly.
修复后的完整代码
Here's the corrected version with fixes and comments explaining each change:
import requests from bs4 import BeautifulSoup import urllib.request import pathlib # Define your root URL properly root_url = 'https://www.ccc.com' class Data: def getlk(self, url): all_links = [] page = requests.get(url) # Add a check to ensure the request succeeded page.raise_for_status() soup = BeautifulSoup(page.text, 'html.parser') for href in soup.find_all(class_='omrlist'): link = href.find('a').get('href') # Make sure we're building absolute URLs correctly if not link.startswith('http'): all_links.append(root_url + link) else: all_links.append(link) return all_links def getmth_lk(self, ylk): rpLk = [] for url in ylk: links = self.getlk(url) # Fix: Extend rpLk with the new links instead of appending lists rpLk.extend(links) for url_ in links: # Filter for non-Annual PDFs if ".pdf" in url_ and "Annual" not in url_: # Extract filename cleanly filename = url_.split('/')[-1] # Use pathlib properly for cross-platform paths download_path = pathlib.Path('/home') / filename if not download_path.exists(): print(f"Downloading {filename}...") # Fix: Pass the actual PDF URL (url_) instead of rpLk urllib.request.urlretrieve(url_, str(download_path)) return rpLk if __name__ == '__main__': obj = Data() # Build the initial URL correctly yLk = obj.getlk(root_url + '/oil/reports/') mth_lk = obj.getmth_lk(yLk) print(f"Processed {len(mth_lk)} total links, downloaded relevant PDFs.")
关键修复说明
- Fixed the URL parameter: Changed
urllib.request.urlretrieve(rpLk, ...)tourllib.request.urlretrieve(url_, ...)to pass the actual PDF link. - Added missing imports: Included all required libraries at the top.
- Defined
root_url: Matched your base URL to avoid broken links. - Fixed list handling: Used
rpLk.extend(links)instead of appending lists to get a flat list of links. - Improved path handling: Used
pathlib.Pathproperly to create cross-platform paths. - Added basic error checking: Used
page.raise_for_status()to catch failed HTTP requests early. - Cleaner filename extraction: Used
split('/')[-1]to get the filename without extra string operations.
内容的提问来源于stack exchange,提问作者P.Jhon
相关产品推荐
相关产品推荐

