You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析大JSON-LD文件:@ID引用合规性及关联方法问询

Is this JSON-LD syntax standard?

Absolutely—this is a core, fully compliant feature of JSON-LD. The @id property exists specifically to uniquely identify nodes in a JSON-LD document, and referencing a node by its @id value (like {"@id": "data:0089c6e8-f891-37b2-90cb-f833741e45ab"}) is the official way to link to that node without repeating its entire data. This pattern is designed to reduce redundancy in large documents, which is exactly what you're dealing with.

Resolving these references with PyLD (no brute force needed)

PyLD was built to handle this kind of internal node linking automatically—you don't need to write custom code to traverse the entire file. Here are two practical approaches depending on your needs:

1. Create a fast lookup map with flatten()

The flatten() function processes your JSON-LD document into a flat array of all nodes, each keyed by their @id. You can then build a dictionary to look up any referenced node in constant time:

import json
from pyld import jsonld

# Load your large JSON-LD file
with open("your_large_file.jsonld", "r") as f:
    json_ld_data = json.load(f)

# Flatten the document to collect all nodes at the top level
flattened_nodes = jsonld.flatten(json_ld_data)

# Build a lookup dictionary for instant access by @id
node_lookup = {node["@id"]: node for node in flattened_nodes}

# Example: Get the full data for your reference
target_id = "data:0089c6e8-f891-37b2-90cb-f833741e45ab"
full_node_data = node_lookup.get(target_id)

if full_node_data:
    print(json.dumps(full_node_data, indent=2))
else:
    print(f"Node with ID {target_id} not found.")

This approach is efficient because PyLD handles the flattening optimally, and the lookup is O(1) once the dictionary is built—no brute-force traversal required.

2. Automatically embed references with frame()

If you want to replace all @id references with their full node data directly in your document structure (so you don't have to look them up separately), use JSON-LD framing. Define a frame that tells PyLD to embed all referenced nodes:

# Define a frame that embeds all referenced nodes
embedding_frame = {
    "@context": json_ld_data["@context"],
    "@embed": "@always"
}

# Apply the frame to your data
framed_data = jsonld.frame(json_ld_data, embedding_frame)

# Now all references are replaced with their full node content
print(json.dumps(framed_data, indent=2))

This will transform your document so that any {"@id": "..."} reference is replaced with the complete node it points to, making it easy to traverse the data as a single, cohesive structure.

Bonus: Handling extremely large files

If your file is too big to load entirely into memory, PyLD supports streaming processing (though it's a bit more involved). You can use the jsonld.stream_expand() method to process the document in chunks while still resolving references. However, for most cases, loading the full document into memory (as shown above) is manageable and simpler.

内容的提问来源于stack exchange,提问作者DJousto

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 10:07:20