JSON文件EOF错误排查及特定关系子集生成求助
JSON Format Explanation
Your train.json uses JSON Lines (newline-delimited JSON) format—each line is a standalone valid JSON object. This is a valid format for batch processing, but standard JSON parsers (like the one in VS Code) expect a single top-level structure (e.g., an array), which is why you get the "EOF expected" error.
If you want to make the file compatible with standard JSON tools, wrap all lines in square brackets and add commas between objects:
[ {"fold":3,"source":"Diffbot Knowledge Graph",...}, {"fold":2,"source":"Diffbot Knowledge Graph",...}, {"fold":1,"source":"Diffbot Knowledge Graph",...} ]
For your Python processing, however, the line-by-line approach is correct and doesn’t require modifying the original file.
Python Code Fix
Your original code tries to access entry['triples'], which doesn’t exist in the JSON data. Relations are stored in two locations:
passages[].exhaustivelyAnnotatedProperties[].propertyName: Directly lists relation names likeEDUCATED_ATpassages[].facts[].propertyId: References a relation ID that maps to a name inexhaustivelyAnnotatedProperties
Here’s the corrected code that properly checks both locations for your target relations:
import json # Define target relations as a set for faster lookups TARGET_RELATIONS = { 'DATE_OF_BIRTH', 'RESIDENCE', 'BIRTHPLACE', 'NATIONALITY', 'EMPLOYEE_OF', 'EDUCATED_AT' } # Load and parse the JSON Lines file data = [] with open('train.json', 'r', encoding="utf8") as f: for line in f: stripped_line = line.strip() if stripped_line: # Skip empty lines to avoid parsing errors data.append(json.loads(stripped_line)) subset = [] for entry in data: has_target_relation = False # Create a map of property IDs to their names for this entry property_id_to_name = {} # First, build the property map and check direct annotations for passage in entry.get('passages', []): for prop in passage.get('exhaustivelyAnnotatedProperties', []): property_id_to_name[prop['propertyId']] = prop['propertyName'] if prop['propertyName'] in TARGET_RELATIONS: has_target_relation = True break if has_target_relation: break # If no direct match found, check facts using the property map if not has_target_relation: for passage in entry.get('passages', []): for fact in passage.get('facts', []): relation_name = property_id_to_name.get(fact['propertyId']) if relation_name in TARGET_RELATIONS: has_target_relation = True break if has_target_relation: break # Add entry to subset if it has a target relation if has_target_relation: subset.append(entry) # Save the subset with readable formatting with open('subset.json', 'w', encoding="utf8") as f: json.dump(subset, f, indent=2)
Key Improvements:
- Efficient relation checks: Uses a set for target relations instead of multiple
orconditions - Robust parsing: Skips empty lines to avoid JSON decoding errors
- Complete relation coverage: Checks both direct property annotations and fact references
- Readable output: Adds indentation to the saved subset JSON
内容的提问来源于stack exchange,提问作者prakash aryan

