You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JSON文件EOF错误排查及特定关系子集生成求助

JSON Format Explanation

Your train.json uses JSON Lines (newline-delimited JSON) format—each line is a standalone valid JSON object. This is a valid format for batch processing, but standard JSON parsers (like the one in VS Code) expect a single top-level structure (e.g., an array), which is why you get the "EOF expected" error.

If you want to make the file compatible with standard JSON tools, wrap all lines in square brackets and add commas between objects:

[
{"fold":3,"source":"Diffbot Knowledge Graph",...},
{"fold":2,"source":"Diffbot Knowledge Graph",...},
{"fold":1,"source":"Diffbot Knowledge Graph",...}
]

For your Python processing, however, the line-by-line approach is correct and doesn’t require modifying the original file.

Python Code Fix

Your original code tries to access entry['triples'], which doesn’t exist in the JSON data. Relations are stored in two locations:

  1. passages[].exhaustivelyAnnotatedProperties[].propertyName: Directly lists relation names like EDUCATED_AT
  2. passages[].facts[].propertyId: References a relation ID that maps to a name in exhaustivelyAnnotatedProperties

Here’s the corrected code that properly checks both locations for your target relations:

import json

# Define target relations as a set for faster lookups
TARGET_RELATIONS = {
    'DATE_OF_BIRTH',
    'RESIDENCE',
    'BIRTHPLACE',
    'NATIONALITY',
    'EMPLOYEE_OF',
    'EDUCATED_AT'
}

# Load and parse the JSON Lines file
data = []
with open('train.json', 'r', encoding="utf8") as f:
    for line in f:
        stripped_line = line.strip()
        if stripped_line:  # Skip empty lines to avoid parsing errors
            data.append(json.loads(stripped_line))

subset = []

for entry in data:
    has_target_relation = False
    # Create a map of property IDs to their names for this entry
    property_id_to_name = {}
    
    # First, build the property map and check direct annotations
    for passage in entry.get('passages', []):
        for prop in passage.get('exhaustivelyAnnotatedProperties', []):
            property_id_to_name[prop['propertyId']] = prop['propertyName']
            if prop['propertyName'] in TARGET_RELATIONS:
                has_target_relation = True
                break
        if has_target_relation:
            break
    
    # If no direct match found, check facts using the property map
    if not has_target_relation:
        for passage in entry.get('passages', []):
            for fact in passage.get('facts', []):
                relation_name = property_id_to_name.get(fact['propertyId'])
                if relation_name in TARGET_RELATIONS:
                    has_target_relation = True
                    break
            if has_target_relation:
                break
    
    # Add entry to subset if it has a target relation
    if has_target_relation:
        subset.append(entry)

# Save the subset with readable formatting
with open('subset.json', 'w', encoding="utf8") as f:
    json.dump(subset, f, indent=2)

Key Improvements:

  • Efficient relation checks: Uses a set for target relations instead of multiple or conditions
  • Robust parsing: Skips empty lines to avoid JSON decoding errors
  • Complete relation coverage: Checks both direct property annotations and fact references
  • Readable output: Adds indentation to the saved subset JSON

内容的提问来源于stack exchange,提问作者prakash aryan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 04:27:01