如何将LangChain生成的Documents转换回Python代码字符串?
将LangChain分割的Documents还原为Python代码
问题分析
你当前的分割函数仅把文件内容拆分为Document对象,但未保留每个chunk对应的原文件信息,也没记录chunk的顺序,导致无法直接还原。解决这个问题需要两步:给每个Document添加原文件元数据,按元数据分组拼接chunk。
修改分割函数(保留元数据)
先修改index_repo函数,确保每个分割后的Document携带原文件路径的元数据,同时修正原代码中可能出现的文件列表与内容列表不匹配的问题:
def index_repo(repo_url): os.environ['OPENAI_API_KEY'] = "" contents = [] fileextensions = [".py"] print('cloning repo') repo_dir = get_repo(repo_url) file_names = [] for dirpath, dirnames, filenames in os.walk(repo_dir): for file in filenames: if file.endswith(tuple(fileextensions)): file_path = os.path.join(dirpath, file) try: with open(file_path, "r", encoding="utf-8") as f: content = f.read() contents.append(content) file_names.append(file_path) except Exception as e: print(f"读取文件失败 {file_path}: {e}") pass # 初始化分割器,同时传入对应文件的元数据 text_splitter = RecursiveCharacterTextSplitter.from_language( language=Language.PYTHON, chunk_size=5000, chunk_overlap=0 ) # 为每个文件创建元数据,关联原文件路径 metadatas = [{"source": filename} for filename in file_names] texts = text_splitter.create_documents(contents, metadatas=metadatas) return texts, file_names
还原为Python代码的函数
编写还原函数,通过元数据将同一文件的chunk分组,再按顺序拼接还原完整代码:
def reconstruct_py_files(documents): # 按原文件路径分组存储chunk内容 file_chunk_map = {} for doc in documents: source_file = doc.metadata.get("source") if not source_file: continue if source_file not in file_chunk_map: file_chunk_map[source_file] = [] file_chunk_map[source_file].append(doc.page_content) # 拼接每个文件的所有chunk,还原完整代码 reconstructed = {} for filename, chunks in file_chunk_map.items(): # 分割器是按文本顺序分割的,直接拼接即可还原原内容 full_code = "".join(chunks) reconstructed[filename] = full_code return reconstructed
使用示例
调用还原函数后,你会得到一个字典,键是原文件路径,值是还原后的完整Python代码:
# 获取分割后的Documents split_docs, _ = index_repo("你的仓库URL") # 还原代码 recovered_files = reconstruct_py_files(split_docs) # 将还原后的代码写入文件 for filename, code in recovered_files.items(): with open(filename, "w", encoding="utf-8") as f: f.write(code)
内容的提问来源于stack exchange,提问作者alpa
相关产品推荐
相关产品推荐

