Python解析大JSON-LD文件:@ID引用合规性及关联方法问询
Absolutely—this is a core, fully compliant feature of JSON-LD. The @id property exists specifically to uniquely identify nodes in a JSON-LD document, and referencing a node by its @id value (like {"@id": "data:0089c6e8-f891-37b2-90cb-f833741e45ab"}) is the official way to link to that node without repeating its entire data. This pattern is designed to reduce redundancy in large documents, which is exactly what you're dealing with.
PyLD was built to handle this kind of internal node linking automatically—you don't need to write custom code to traverse the entire file. Here are two practical approaches depending on your needs:
1. Create a fast lookup map with flatten()
The flatten() function processes your JSON-LD document into a flat array of all nodes, each keyed by their @id. You can then build a dictionary to look up any referenced node in constant time:
import json from pyld import jsonld # Load your large JSON-LD file with open("your_large_file.jsonld", "r") as f: json_ld_data = json.load(f) # Flatten the document to collect all nodes at the top level flattened_nodes = jsonld.flatten(json_ld_data) # Build a lookup dictionary for instant access by @id node_lookup = {node["@id"]: node for node in flattened_nodes} # Example: Get the full data for your reference target_id = "data:0089c6e8-f891-37b2-90cb-f833741e45ab" full_node_data = node_lookup.get(target_id) if full_node_data: print(json.dumps(full_node_data, indent=2)) else: print(f"Node with ID {target_id} not found.")
This approach is efficient because PyLD handles the flattening optimally, and the lookup is O(1) once the dictionary is built—no brute-force traversal required.
2. Automatically embed references with frame()
If you want to replace all @id references with their full node data directly in your document structure (so you don't have to look them up separately), use JSON-LD framing. Define a frame that tells PyLD to embed all referenced nodes:
# Define a frame that embeds all referenced nodes embedding_frame = { "@context": json_ld_data["@context"], "@embed": "@always" } # Apply the frame to your data framed_data = jsonld.frame(json_ld_data, embedding_frame) # Now all references are replaced with their full node content print(json.dumps(framed_data, indent=2))
This will transform your document so that any {"@id": "..."} reference is replaced with the complete node it points to, making it easy to traverse the data as a single, cohesive structure.
Bonus: Handling extremely large files
If your file is too big to load entirely into memory, PyLD supports streaming processing (though it's a bit more involved). You can use the jsonld.stream_expand() method to process the document in chunks while still resolving references. However, for most cases, loading the full document into memory (as shown above) is manageable and simpler.
内容的提问来源于stack exchange,提问作者DJousto

