Yelp数据集JSON转CSV脚本无输出问题排查求助(Python3)
Let's walk through the issues with your script and get it generating CSV files properly, plus answer your questions about file paths and parameters.
1. The Critical Reason No CSV is Being Generated
Your script's main execution block isn't running at all! You have a typo in the entry point check:
if name == 'main': # Wrong
This should be Python's standard double-underscore syntax:
if __name__ == '__main__': # Correct
Without the underscores, the code that parses arguments, generates the CSV filename, and runs the conversion never executes—so the script runs but does nothing.
2. Other Script Syntax & Logic Fixes
Broken Import Statement
Your first line tries to import multiple modules in one line, which is invalid. Split them into separate lines:
import argparse import collections import csv import json
Encoding Consistency
In get_superset_of_column_names_from_file, you open the JSON file without specifying encoding="utf8"—this can cause encoding errors on Windows. Update that function:
def get_superset_of_column_names_from_file(json_file_path): column_names = set() with open(json_file_path, encoding="utf8") as fin: # Added encoding for line in fin: line_contents = json.loads(line) column_names.update( set(get_column_names(line_contents).keys()) ) return column_names
Clean Up String Handling in CSV Rows
Your get_row function encodes strings to bytes, which will write b'your text' into the CSV (not what you want). Replace that logic to write raw strings (and handle newlines to avoid CSV formatting issues):
def get_row(line_contents, column_names): row = [] for column_name in column_names: line_value = get_nested_value(line_contents, column_name) if isinstance(line_value, str): # Replace newlines to prevent breaking CSV rows row.append(line_value.replace('\n', ' ').replace('\r', '')) elif line_value is not None: row.append(str(line_value)) else: row.append('') return row
3. File Path & Parameter Questions
Current Behavior (No Manual Path Needed)
Right now, your script:
- Takes the input JSON filename as a positional argument
- Automatically generates the CSV filename by replacing
.jsonwith.csv(e.g.,review.json→review.csv) - Saves the CSV in your current working directory (the folder you're in when running the command)
How to Manually Specify Output Path/Name
If you want to control where the CSV is saved, add an optional argument to the parser. Update the __main__ block like this:
if __name__ == '__main__': parser = argparse.ArgumentParser( description='Convert Yelp data from JSON format to CSV.', ) parser.add_argument( 'json_file', type=str, help='The json file to convert (can be a full path or relative path).', ) # Add optional output argument parser.add_argument( '--output', '-o', type=str, help='Optional: Full or relative path for the output CSV file. If not provided, uses input filename with .csv extension.', default=None ) args = parser.parse_args() json_file = args.json_file # Use specified output path if provided, else auto-generate if args.output: csv_file = args.output else: csv_file = '{0}.csv'.format(json_file.split('.json')[0]) column_names = get_superset_of_column_names_from_file(json_file) read_and_write_file(json_file, csv_file, column_names)
Now you can run it with a custom output path like:
C:\Users\AppData\Local\Programs\Python\Python36-32>python.exe json_to_csv_converter.py review.json -o C:\MyData\yelp_reviews.csv
Correct Command Execution Tips
- If your
review.jsonisn't in the same folder as your script (or the folder you're running the command from), use the full path to the JSON file:python.exe json_to_csv_converter.py C:\Path\To\Your\Data\review.json - Alternatively, navigate to the JSON file's folder first, then run the script with the relative path:
cd C:\Path\To\Your\Data C:\Users\AppData\Local\Programs\Python\Python36-32\python.exe C:\Path\To\Your\Script\json_to_csv_converter.py review.json
Full Corrected Script
import argparse import collections import csv import json def read_and_write_file(json_file_path, csv_file_path, column_names): with open(csv_file_path, 'w+', newline='', encoding="utf8") as fout: # Added newline='' and encoding csv_file = csv.writer(fout) csv_file.writerow(list(column_names)) with open(json_file_path, encoding="utf8") as fin: for line in fin: line_contents = json.loads(line) csv_file.writerow(get_row(line_contents, column_names)) def get_superset_of_column_names_from_file(json_file_path): column_names = set() with open(json_file_path, encoding="utf8") as fin: for line in fin: line_contents = json.loads(line) column_names.update( set(get_column_names(line_contents).keys()) ) return column_names def get_column_names(line_contents, parent_key=''): column_names = [] for k, v in line_contents.items(): column_name = "{0}.{1}".format(parent_key, k) if parent_key else k if isinstance(v, collections.MutableMapping): column_names.extend( get_column_names(v, column_name).items() ) else: column_names.append((column_name, v)) return dict(column_names) def get_nested_value(d, key): if '.' not in key: if key not in d: return None return d[key] base_key, sub_key = key.split('.', 1) if base_key not in d: return None sub_dict = d[base_key] return get_nested_value(sub_dict, sub_key) def get_row(line_contents, column_names): row = [] for column_name in column_names: line_value = get_nested_value(line_contents, column_name) if isinstance(line_value, str): row.append(line_value.replace('\n', ' ').replace('\r', '')) elif line_value is not None: row.append(str(line_value)) else: row.append('') return row if __name__ == '__main__': parser = argparse.ArgumentParser( description='Convert Yelp data from JSON format to CSV.', ) parser.add_argument( 'json_file', type=str, help='The json file to convert (can be a full path or relative path).', ) parser.add_argument( '--output', '-o', type=str, help='Optional: Full or relative path for the output CSV file. If not provided, uses input filename with .csv extension.', default=None ) args = parser.parse_args() json_file = args.json_file if args.output: csv_file = args.output else: csv_file = '{0}.csv'.format(json_file.split('.json')[0]) column_names = get_superset_of_column_names_from_file(json_file) read_and_write_file(json_file, csv_file, column_names)
内容的提问来源于stack exchange,提问作者carmem

