You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用BeautifulSoup4爬取GitHub时出现文本缩进异常问题求助

问题原因

你遇到的输出缩进异常是因为从GitHub页面提取的<p>标签文本本身自带了前后换行、多余空白字符,属于前端页面原始格式残留,和你使用的replit运行环境无关。

解决方法

直接使用BeautifulSoup内置的get_text()方法,传入strip=True参数即可自动清理文本前后所有多余的空白、换行符,还会自动合并文本内部的连续空白。修正后可运行的完整代码如下:

import requests as req
from bs4 import BeautifulSoup

# 可选:添加请求头伪装浏览器,避免被GitHub反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}

urls = [
  "https://github.com/moom825/Discord-RAT",
  "https://github.com/freyacodes/Lavalink",
  "https://github.com/KagChi/lavalink-railways",
  "https://github.com/KagChi/lavalink-repl",
  "https://github.com/Devoxin/Lavalink.py",
  "https://github.com/karyeet/heroku-lavalink"]

# 添加timeout参数防止请求无响应卡住
r = req.get(urls[0], headers=headers, timeout=10)
soup = BeautifulSoup(r.content,"lxml")

desc_elem = soup.find("p",attrs={"class":"f4 mt-3"})
# 先判断元素是否存在,避免无仓库描述时报错
if desc_elem:
    # 使用get_text(strip=True)清理多余格式
    title = desc_elem.get_text(strip=True)
    print(title)
else:
    print("未找到仓库描述信息")

额外说明

  • 如果需要保留文本内部的换行,仅清理前后空白,可以将strip=True换成手动调用str.strip()方法,即title = desc_elem.text.strip()
  • 批量爬取多个GitHub仓库时建议添加请求间隔,避免触发GitHub的反爬机制被临时限制访问。

内容的提问来源于stack exchange,提问作者DevER-M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 00:15:02