Python实现IPLD Merkle DAG节点链接及CAR文件生成技术求助
Merkle DAG节点链接规范与CAR文件生成问题
我正在开发Python实现的Merkle DAG,目标是生成Content Addressable Archive(CAR)文件,但在节点链接的规范实现上遇到瓶颈。使用multiformats库为数据块生成CID,尝试将子块的CID存储在根节点的"links"字段中,但不确定是否符合IPLD Merkle DAG的节点链接要求。以下是我的Python3实现代码:
from multiformats import CID, varint, multihash, multibase import dag_cbor import json import msgpack def generate_cid(data, codec="dag-pb"): hash_value = multihash.digest(data, "sha2-256") return CID("base32", version=1, codec=codec, digest=hash_value) def generate_merkle_tree(file_path, chunk_size): cids = [] # Read the file with open(file_path, "rb") as file: while True: # Read a chunk of data chunk = file.read(chunk_size) if not chunk: break # Generate CID for the chunk cid = generate_cid(chunk, codec="raw") cids.append((cid, chunk)) # Generate Merkle tree root CID from all the chunks # root_cid = generate_cid(b"".join(bytes(cid[0]) for cid in cids)) # Create the root node with links and other data root_node = { "file_name": "test.png", "links": [str(cid[0]) for cid in cids] } # Encode the root node as dag-pb root_data = dag_cbor.encode(root_node) # Generate CID for the root node root_cid = generate_cid(root_data, codec="dag-pb") return root_cid, cids, root_data def create_car_file(root, cids): header_roots = [root] header_data = dag_cbor.encode({"roots": header_roots, "version": 1}) header = varint.encode(len(header_data)) + header_data car_content = b"" car_content += header for cid, chunk in cids: cid_bytes = bytes(cid) block = varint.encode(len(chunk) + len(cid_bytes)) + cid_bytes + chunk car_content += block root_cid = bytes(root) root_block = varint.encode(len(root_cid)) + root_cid car_content += root_block with open("output.car", "wb") as car_file: car_file.write(car_content) file_path = "./AADHAAR.png" # Replace with the path to your file chunk_size = 16384 # Adjust the chunk size as needed root, cids, root_data = generate_merkle_tree(file_path, chunk_size) print(root) create_car_file(root, cids)
代码审核与修正建议
1. 核心问题:节点链接与编码不匹配
当前实现混淆了IPLD的dag-pb和dag-cbor编码规范,导致节点无法被IPLD解析:
- 编码与codec不一致:用
dag_cbor.encode生成根节点数据,却指定codec为dag-pb,这会导致CID的编码标识与实际数据格式不匹配。 - links字段格式错误:IPLD规范中,links不是字符串数组,而是包含CID、元数据的结构化对象,且需存储CID对象而非字符串。
修正方案(二选一)
方案A:使用dag-cbor编码(灵活自定义结构)
根节点可以是任意CBOR可序列化结构,links存储CID对象即可:
def generate_merkle_tree(file_path, chunk_size): cids = [] with open(file_path, "rb") as file: while True: chunk = file.read(chunk_size) if not chunk: break cid = generate_cid(chunk, codec="raw") cids.append((cid, chunk)) # 构造符合dag-cbor规范的根节点 root_node = { "file_name": "test.png", "links": [ {"cid": cid, "size": len(chunk), "name": f"chunk-{i}"} for i, (cid, chunk) in enumerate(cids) ] } root_data = dag_cbor.encode(root_node) # 指定正确的codec:dag-cbor root_cid = generate_cid(root_data, codec="dag-cbor") return root_cid, cids, root_data
方案B:使用dag-pb编码(IPFS兼容标准格式)
需安装ipld-dag-pb库,遵循严格的protobuf节点结构:
from ipld_dag_pb import PBNode, PBLink def generate_merkle_tree(file_path, chunk_size): cids = [] with open(file_path, "rb") as file: while True: chunk = file.read(chunk_size) if not chunk: break cid = generate_cid(chunk, codec="raw") cids.append((cid, chunk)) # 构造dag-pb规范的节点链接 pb_links = [] for i, (cid, chunk) in enumerate(cids): pb_link = PBLink(hash=cid, name=f"chunk-{i}", size=len(chunk)) pb_links.append(pb_link) # dag-pb节点的Data字段存储自定义元数据(如文件名) pb_node = PBNode(data=b"test.png", links=pb_links) root_data = pb_node.SerializeToString() root_cid = generate_cid(root_data, codec="dag-pb") return root_cid, cids, root_data
2. CAR文件生成错误修正
当前create_car_file函数存在两个致命问题,导致CAR文件无效:
- 根节点块仅写入了CID字节,未写入实际节点数据;
- 头部
roots字段需存储CID的二进制形式,而非CID对象。
修正后的函数:
def create_car_file(root_cid, cids, root_data): # 构造CAR v1头部,roots为CID二进制数组 header = { "version": 1, "roots": [bytes(root_cid)] } header_data = dag_cbor.encode(header) car_content = varint.encode(len(header_data)) + header_data # 写入所有数据块 for cid, chunk in cids: cid_bytes = bytes(cid) block_len = len(cid_bytes) + len(chunk) car_content += varint.encode(block_len) + cid_bytes + chunk # 写入根节点块 root_cid_bytes = bytes(root_cid) root_block_len = len(root_cid_bytes) + len(root_data) car_content += varint.encode(root_block_len) + root_cid_bytes + root_data with open("output.car", "wb") as car_file: car_file.write(car_content)
3. 其他优化点
- 简化CID生成函数:去掉冗余的
base32参数,multiformats默认会处理编码:def generate_cid(data, codec="dag-cbor"): mh = multihash.digest(data, "sha2-256") return CID(1, codec, mh) - 验证CAR文件:生成后可通过
ipfs car verify output.car命令验证有效性(需安装IPFS)。
内容的提问来源于stack exchange,提问作者Ritesh Saha
相关产品推荐
相关产品推荐

