You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3下载PDF时报TypeError: expected string or bytes-like object错误求助

解决Python下载PDF时的TypeError问题

Hey there, let's break down what's going wrong here and fix it step by step!

直接错误原因

The TypeError you're seeing comes down to one critical mistake in your getmth_lk method:
You're passing rpLk (an empty list you initialized) to urllib.request.urlretrieve() instead of the actual PDF URL (url_) from your loop. The urlretrieve() function expects a string/bytes URL, not a list—hence the error.

其他需要修复的问题

Beyond that, there are a few other issues in your code that will cause errors or unexpected behavior:

  • Missing imports: You're using requests, BeautifulSoup, urllib, and pathlib but haven't imported them.
  • Undefined variables: reportLinks and root_url aren't defined anywhere in your code.
  • Incorrect list appending: You're trying to append lists directly to rpLk instead of extending it (or using it correctly).
  • Path handling: Hardcoding /home/ might not be portable, and you're not using the path object you created properly.

修复后的完整代码

Here's the corrected version with fixes and comments explaining each change:

import requests
from bs4 import BeautifulSoup
import urllib.request
import pathlib

# Define your root URL properly
root_url = 'https://www.ccc.com'

class Data:
    def getlk(self, url):
        all_links = []
        page = requests.get(url)
        # Add a check to ensure the request succeeded
        page.raise_for_status()
        soup = BeautifulSoup(page.text, 'html.parser')
        for href in soup.find_all(class_='omrlist'):
            link = href.find('a').get('href')
            # Make sure we're building absolute URLs correctly
            if not link.startswith('http'):
                all_links.append(root_url + link)
            else:
                all_links.append(link)
        return all_links
    
    def getmth_lk(self, ylk):
        rpLk = []
        for url in ylk:
            links = self.getlk(url)
            # Fix: Extend rpLk with the new links instead of appending lists
            rpLk.extend(links)
            for url_ in links:
                # Filter for non-Annual PDFs
                if ".pdf" in url_ and "Annual" not in url_:
                    # Extract filename cleanly
                    filename = url_.split('/')[-1]
                    # Use pathlib properly for cross-platform paths
                    download_path = pathlib.Path('/home') / filename
                    if not download_path.exists():
                        print(f"Downloading {filename}...")
                        # Fix: Pass the actual PDF URL (url_) instead of rpLk
                        urllib.request.urlretrieve(url_, str(download_path))
        return rpLk

if __name__ == '__main__':
    obj = Data()
    # Build the initial URL correctly
    yLk = obj.getlk(root_url + '/oil/reports/')
    mth_lk = obj.getmth_lk(yLk)
    print(f"Processed {len(mth_lk)} total links, downloaded relevant PDFs.")

关键修复说明

  • Fixed the URL parameter: Changed urllib.request.urlretrieve(rpLk, ...) to urllib.request.urlretrieve(url_, ...) to pass the actual PDF link.
  • Added missing imports: Included all required libraries at the top.
  • Defined root_url: Matched your base URL to avoid broken links.
  • Fixed list handling: Used rpLk.extend(links) instead of appending lists to get a flat list of links.
  • Improved path handling: Used pathlib.Path properly to create cross-platform paths.
  • Added basic error checking: Used page.raise_for_status() to catch failed HTTP requests early.
  • Cleaner filename extraction: Used split('/')[-1] to get the filename without extra string operations.

内容的提问来源于stack exchange,提问作者P.Jhon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:35:58