You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Telegra.ph文章抓取合并PDF时出现requests连接超时错误求助

解决抓取telegra.ph文章时的ConnectionError超时问题

你遇到的requests.exceptions.ConnectionError(错误码10060)属于网络连接超时问题,大概率是以下原因导致:

  • 本地网络无法直接访问telegra.ph平台
  • 请求未设置超时时间,导致连接长时间挂起后被终止
  • 目标服务器对请求频率做了限制

针对这些问题,结合你的需求,给出以下修复方案和优化后的代码:

核心修复点

1. 添加请求超时与异常捕获

给requests.get添加超时参数,同时捕获连接异常,避免程序直接崩溃,还能输出具体错误信息。

2. 可选:配置代理访问

如果是本地网络限制导致无法访问telegra.ph,可添加代理配置(根据自己的代理信息修改)。

3. 优化PDF文本排版

原代码中直接按段落换行绘制文本,当内容过长时会超出页面范围,这里优化为自动换行处理。

修改后的完整代码

import requests
from bs4 import BeautifulSoup
from reportlab.pdfgen import canvas
from reportlab.lib.pagesizes import letter
from reportlab.lib.utils import simpleSplit

def fetch_telegraph_content(url, proxy=None):
    try:
        # 设置10秒超时,避免无限等待
        response = requests.get(url, timeout=10, proxies=proxy)
        response.raise_for_status()  # 主动抛出HTTP错误
        soup = BeautifulSoup(response.text, 'html.parser')
        title = soup.find('title').text
        paragraphs = soup.find_all('p')
        # 合并段落,用空行分隔
        content = '\n\n'.join([p.text.strip() for p in paragraphs if p.text.strip()])
        return title, content
    except requests.exceptions.RequestException as e:
        print(f"抓取{url}失败: {str(e)}")
        return None, None

def create_pdf(filename, articles):
    c = canvas.Canvas(filename, pagesize=letter)
    width, height = letter
    margin = 72  # 页面边距
    line_height = 14  # 行高
    current_y = height - margin  # 初始Y坐标

    for title, content in articles:
        # 绘制标题
        c.setFont("Helvetica-Bold", 14)
        c.drawString(margin, current_y, title)
        current_y -= 30  # 标题下方留空

        # 绘制内容,自动换行
        c.setFont("Helvetica", 12)
        # 按页面宽度分割每行文本
        lines = simpleSplit(content, "Helvetica", 12, width - 2*margin)
        for line in lines:
            if current_y < margin:  # 内容超出当前页,新建页面
                c.showPage()
                current_y = height - margin
                c.setFont("Helvetica", 12)
            c.drawString(margin, current_y, line)
            current_y -= line_height
        
        c.showPage()  # 每篇文章结束后新建页面
        current_y = height - margin  # 重置Y坐标

    c.save()

def main():
    urls = [
        "https://telegra.ph/1234",
        "https://telegra.ph/123"
    ]
    # 可选:如果需要代理,取消注释并修改代理地址
    # proxy = {
    #     'http': 'http://your-proxy:port',
    #     'https': 'https://your-proxy:port'
    # }
    proxy = None

    articles = []
    for url in urls:
        title, content = fetch_telegraph_content(url, proxy)
        if title and content:
            articles.append((title, content))

    if articles:
        create_pdf("Telegraph_Articles.pdf", articles)
        print("PDF生成成功: Telegraph_Articles.pdf")
    else:
        print("未抓取到任何文章,无法生成PDF。")

if __name__ == "__main__":
    main()

额外说明

  • 超时时间可根据网络情况调整,建议设置5-20秒
  • 代理配置需根据实际可用的代理服务填写,若无需代理则保持proxy=None
  • 优化后的PDF生成逻辑会自动处理长文本换行和分页,避免内容被截断

内容的提问来源于stack exchange,提问作者Adilkhan Dilman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 10:33:17