如何爬取GitHub仓库代码?现有Python代码返回空响应
解决GitHub仓库代码爬取问题
你的代码出现空响应和无法提取函数的问题,核心原因有几个,下面给出具体修正方案:
问题分析
- 请求URL错误:你请求的是GitHub的blob展示页面,该页面代码块结构复杂且可能包含动态加载内容,远不如直接请求原始代码文件可靠。
- 函数名匹配错误:你定义的
functions_to_fetch里是带()的函数名,但GitHub代码高亮中,nf类的span仅显示函数名(不带()),导致匹配失败。 - HTML解析逻辑错误:判断
def的代码逻辑有误,code_block.find('span', {'class': 'kd'})返回的是元素对象,直接和字符串'def'比较永远不成立,需要取其文本内容再判断。
修正后的代码
import requests # 定义目标仓库和文件路径 repo_owner = "Arbaxali" repo_name = "Memepage-using-flask" file_path = "memepage.py" # 要提取的函数名(不带()) functions_to_fetch = ["get_meme", "index"] # 请求原始代码文件,而非blob展示页面 url = f"https://github.com/{repo_owner}/{repo_name}/raw/master/{file_path}" response = requests.get(url) if response.status_code == 200: # 获取原始代码文本并按行拆分 code_lines = response.text.splitlines() functions = {} current_func = None current_code = [] for line in code_lines: stripped_line = line.strip() # 判断是否为函数定义行 if stripped_line.startswith('def ') and stripped_line.endswith(':'): func_name = stripped_line.split('def ')[1].split('(')[0] # 匹配到目标函数则开始收集代码 if func_name in functions_to_fetch: current_func = func_name current_code.append(line) else: current_func = None elif current_func is not None: # 收集函数内部代码,直到遇到新函数或非缩进内容 if stripped_line and not stripped_line.startswith('def '): current_code.append(line) else: functions[current_func] = '\n'.join(current_code) current_func = None current_code = [] # 处理最后一个未闭合的函数 if current_func is not None: functions[current_func] = '\n'.join(current_code) # 打印提取结果 for func_name, func_code in functions.items(): print(f"函数名: {func_name}\n函数代码:\n{func_code}\n") else: print(f"请求失败,状态码: {response.status_code}")
关键改进点
- 改用raw链接:直接获取原始代码文本,无需解析复杂HTML结构,效率和稳定性更高。
- 修正函数名匹配:去掉函数名后的
(),确保匹配逻辑正确。 - 基于文本解析函数:直接对代码文本进行行处理,避免依赖GitHub的HTML结构变动,适配性更强。
内容的提问来源于stack exchange,提问作者Arbaz ali
相关产品推荐
相关产品推荐

