如何编写类预处理并存储LA Times半结构化格式的文章集合?
Alright, let's build a robust class to handle your LA Times article processing task. Below is a complete implementation with explanations for each step.
LATimes Article Processor Class
Full Implementation
import re import json from typing import List, Dict class LATimesArticleProcessor: def __init__(self, file_path: str): self.file_path = file_path self.articles: List[Dict] = [] def load_and_extract_articles(self) -> None: """Load the text file and extract individual articles using <doc> tags""" with open(self.file_path, 'r', encoding='utf-8') as f: content = f.read() # Regex to capture all <doc> blocks (handles multi-line content) doc_pattern = re.compile(r'<doc>(.*?)</doc>', re.DOTALL) doc_blocks = doc_pattern.findall(content) for block in doc_blocks: # Extract article ID (matches lines like "id: 1234") id_match = re.search(r'id:\s*(\d+)', block) article_id = id_match.group(1) if id_match else None # Extract article title (matches lines like "title: Some News Story") title_match = re.search(r'title:\s*(.*)', block) title = title_match.group(1).strip() if title_match else None # Extract raw text between <text> tags text_pattern = re.compile(r'<text>(.*?)</text>', re.DOTALL) text_match = text_pattern.search(block) raw_text = text_match.group(1).strip() if text_match else None # Only add the article if all required fields are present if all([article_id, title, raw_text]): self.articles.append({ 'id': article_id, 'title': title, 'raw_text': raw_text }) def preprocess_text(self, text: str) -> str: """Basic text cleaning pipeline - customize based on your needs""" # Remove extra newlines and redundant spaces cleaned_text = re.sub(r'\s+', ' ', text) # Convert to lowercase (optional, skip if case matters for your use case) cleaned_text = cleaned_text.lower() # Remove special characters (keep letters, numbers, and spaces) cleaned_text = re.sub(r'[^a-zA-Z0-9\s]', '', cleaned_text) return cleaned_text def process_all_articles(self) -> None: """Run the full extraction + preprocessing pipeline""" self.load_and_extract_articles() for article in self.articles: article['processed_text'] = self.preprocess_text(article['raw_text']) def save_to_json(self, output_path: str) -> None: """Store processed articles in a structured JSON file""" with open(output_path, 'w', encoding='utf-8') as f: json.dump(self.articles, f, indent=2) # Example usage if __name__ == "__main__": # Initialize processor with your input file processor = LATimesArticleProcessor('data_science_assignment.txt') # Run full processing pipeline processor.process_all_articles() # Save results to JSON processor.save_to_json('processed_latimes_articles.json')
Key Method Breakdown
load_and_extract_articles(): Uses regex to pull out each<doc>block, then parses the ID, title, and raw text from within. There.DOTALLflag ensures we correctly capture multi-line content inside the tags.preprocess_text(): A baseline cleaning function—you can extend this with stopword removal, lemmatization (using libraries like NLTK or spaCy), or custom rules based on your specific data needs.process_all_articles(): Orchestrates the full workflow: first extracts all valid articles, then applies preprocessing to each one.save_to_json(): Stores the processed data in a human-readable, structured format that's easy to use for downstream tasks like data analysis or NLP modeling.
Customization Ideas
- Swap out
save_to_jsonfor a CSV writer if you prefer tabular format. - Add error handling (try-except blocks) to gracefully handle malformed entries in the input file.
- Integrate advanced NLP tools (like spaCy) for more sophisticated preprocessing (e.g., named entity recognition, dependency parsing).
内容的提问来源于stack exchange,提问作者Samuel Oyediran
相关产品推荐
相关产品推荐

