基于布尔表达式查询含指定关键词的JSON文件名
Got it, let's tackle this problem step by step. The goal is to filter JSON files based on boolean queries against their keyword fields, right? Here's a practical, Python-based solution that handles all the cases you mentioned, including & (AND), ! (NOT), | (OR), and even nested combinations.
核心思路
We need to do two key things:
- Extract the full set of keywords from each JSON file's
keywordfield. - Parse the user's boolean query into executable logic, then check which files match the query.
1. Extract Keywords from JSON Files
First, let's write a helper function to read each JSON file and pull out its keywords. We'll handle both array-style keywords (like "keyword": ["Alabama", "Washington"]) and comma-separated string formats (in case your JSON has typos like the example you shared).
import json from pathlib import Path def extract_keywords(json_file): try: with open(json_file, 'r') as f: data = json.load(f) raw_keywords = data.get('keyword', []) # Handle comma-separated strings (if your JSON uses this format) if isinstance(raw_keywords, str): keywords = {kw.strip() for kw in raw_keywords.split(',')} else: # Assume it's an array/list of keywords keywords = set(raw_keywords) return keywords except Exception as e: print(f"⚠️ Error processing {json_file}: {str(e)}") return set()
2. Parse & Execute Boolean Queries
Next, we need to convert the user's query string (like "Washington & !Alabama") into valid Python logic. We'll replace the query's syntax with Python's built-in logical operators and evaluate the condition against each file's keyword set.
import re def matches_query(query, file_keywords): # Extract all quoted keywords from the query quoted_keywords = re.findall(r'"([^"]+)"', query) # Replace each quoted keyword with a check for membership in the file's keywords transformed_query = query for kw in quoted_keywords: transformed_query = transformed_query.replace(f'"{kw}"', f'"{kw}" in file_keywords') # Map custom boolean operators to Python's syntax transformed_query = ( transformed_query .replace('&', 'and') .replace('|', 'or') .replace('!', 'not ') ) # Safely evaluate the query (only safe if you control the input queries!) try: return eval(transformed_query) except SyntaxError: print(f"❌ Invalid query syntax: {query}") return False
3. Main Script: Tie It All Together
Finally, a main function that handles command-line input, iterates over the specified files, and prints the ones that match the query.
def main(): import sys if len(sys.argv) < 3: print("Usage: python query.py <boolean_query> <file1.json> <file2.json> ...") print("Example: python query.py \"\\\"Washington\\\" & !\\\"Alabama\\\"\" 02.json") sys.exit(1) query = sys.argv[1] file_paths = sys.argv[2:] for file_path in file_paths: path = Path(file_path) if not path.exists(): print(f"⚠️ File {file_path} not found") continue if path.suffix != '.json': print(f"⚠️ {file_path} is not a JSON file") continue file_keywords = extract_keywords(path) if matches_query(query, file_keywords): print(path.name) if __name__ == "__main__": main()
Usage Examples
Let's test this with your sample cases:
1. Match files containing "Washington"
python query.py "\"Washington\"" 01.json 02.json
Output:
01.json 02.json
2. Match files with "Washington" but NOT "Alabama"
python query.py "\"Washington\" & !\"Alabama\"" 02.json
Output:
02.json
3. Match files with "Alabama" OR "Pennsylvania"
python query.py "\"Alabama\" | \"Pennsylvania\"" 01.json 02.json
Output:
01.json 02.json
Important Notes
- Command-line Quoting: In bash/zsh, you need to escape inner quotes (like
\") or wrap the entire query in single quotes (e.g.,' "Washington" & !"Alabama" '). - JSON Format: The code handles both array and comma-separated string
keywordfields. If your JSON uses another format, adjust theextract_keywordsfunction accordingly. - Query Safety: Using
eval()is safe here only if you trust the input queries. If this script will handle untrusted user input, you'll need a proper boolean expression parser instead (like usingpyparsingto build an abstract syntax tree).
内容的提问来源于stack exchange,提问作者nightcod3r

