You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:能否用os.path.join()拼接URL?代码TypeError报错求解决

解决你的Python爬虫报错与URL拼接疑问

一、先搞定TypeError报错

你遇到的TypeError: join() argument must be str or bytes, not 'int'原因很明确:

  • 你的page变量是整数类型,而os.path.join()要求所有传入的路径组件必须是字符串或字节。在代码第9行的os.path.join(storing_folder, chapter, page)里,page是int类型,把它转成字符串str(page)就能解决这个问题。
  • 另外,你第10行的imageFile.write()是空调用,应该写入请求到的图片内容res.content,不然会导致文件写入失败。

修正后的这部分代码应该是这样:

imageFile = open(os.path.join(storing_folder, chapter, str(page)), 'wb')
imageFile.write(res.content)
imageFile.close()

二、关于URL拼接:别用os.path.join()!

os.path.join()是为本地文件系统路径设计的工具,它会根据操作系统自动调整路径分隔符(比如Windows用\,macOS/Linux用/),但URL的分隔符固定是/,用它拼接URL很容易出问题(比如在Windows环境下生成带\的非法URL)。

更优的URL拼接方式推荐这两种:

1. 用f-string直接格式化(最直观)

比如构造章节和页面的URL:

chapter_num = ch_numb_regex.search(chapter).group()
page_url = f"{starting_url}/{chapter_num}/{page}"
res = requests.get(page_url)

2. 用标准库的urllib.parse.urljoin(更严谨)

这个工具能正确处理URL的相对路径和绝对路径,避免格式错误:

from urllib.parse import urljoin

chapter_num = ch_numb_regex.search(chapter).group()
base_chapter_url = urljoin(starting_url, chapter_num)
page_url = urljoin(base_chapter_url, str(page))
res = requests.get(page_url)

额外小建议

  • 可以用with语句来管理文件写入,这样不用手动调用close(),更安全且代码更简洁:
with open(os.path.join(storing_folder, chapter, str(page)), 'wb') as imageFile:
    imageFile.write(res.content)
  • 循环里的while True建议加个终止条件,比如当返回的页面找不到图片时就退出,避免出现无限循环的情况。

内容的提问来源于stack exchange,提问作者ulibrau

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:58:18