You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写类预处理并存储LA Times半结构化格式的文章集合?

Alright, let's build a robust class to handle your LA Times article processing task. Below is a complete implementation with explanations for each step.

LATimes Article Processor Class

Full Implementation

import re
import json
from typing import List, Dict

class LATimesArticleProcessor:
    def __init__(self, file_path: str):
        self.file_path = file_path
        self.articles: List[Dict] = []

    def load_and_extract_articles(self) -> None:
        """Load the text file and extract individual articles using <doc> tags"""
        with open(self.file_path, 'r', encoding='utf-8') as f:
            content = f.read()
        
        # Regex to capture all <doc> blocks (handles multi-line content)
        doc_pattern = re.compile(r'<doc>(.*?)</doc>', re.DOTALL)
        doc_blocks = doc_pattern.findall(content)

        for block in doc_blocks:
            # Extract article ID (matches lines like "id: 1234")
            id_match = re.search(r'id:\s*(\d+)', block)
            article_id = id_match.group(1) if id_match else None

            # Extract article title (matches lines like "title: Some News Story")
            title_match = re.search(r'title:\s*(.*)', block)
            title = title_match.group(1).strip() if title_match else None

            # Extract raw text between <text> tags
            text_pattern = re.compile(r'<text>(.*?)</text>', re.DOTALL)
            text_match = text_pattern.search(block)
            raw_text = text_match.group(1).strip() if text_match else None

            # Only add the article if all required fields are present
            if all([article_id, title, raw_text]):
                self.articles.append({
                    'id': article_id,
                    'title': title,
                    'raw_text': raw_text
                })

    def preprocess_text(self, text: str) -> str:
        """Basic text cleaning pipeline - customize based on your needs"""
        # Remove extra newlines and redundant spaces
        cleaned_text = re.sub(r'\s+', ' ', text)
        # Convert to lowercase (optional, skip if case matters for your use case)
        cleaned_text = cleaned_text.lower()
        # Remove special characters (keep letters, numbers, and spaces)
        cleaned_text = re.sub(r'[^a-zA-Z0-9\s]', '', cleaned_text)
        return cleaned_text

    def process_all_articles(self) -> None:
        """Run the full extraction + preprocessing pipeline"""
        self.load_and_extract_articles()
        for article in self.articles:
            article['processed_text'] = self.preprocess_text(article['raw_text'])

    def save_to_json(self, output_path: str) -> None:
        """Store processed articles in a structured JSON file"""
        with open(output_path, 'w', encoding='utf-8') as f:
            json.dump(self.articles, f, indent=2)

# Example usage
if __name__ == "__main__":
    # Initialize processor with your input file
    processor = LATimesArticleProcessor('data_science_assignment.txt')
    # Run full processing pipeline
    processor.process_all_articles()
    # Save results to JSON
    processor.save_to_json('processed_latimes_articles.json')

Key Method Breakdown

  • load_and_extract_articles(): Uses regex to pull out each <doc> block, then parses the ID, title, and raw text from within. The re.DOTALL flag ensures we correctly capture multi-line content inside the tags.
  • preprocess_text(): A baseline cleaning function—you can extend this with stopword removal, lemmatization (using libraries like NLTK or spaCy), or custom rules based on your specific data needs.
  • process_all_articles(): Orchestrates the full workflow: first extracts all valid articles, then applies preprocessing to each one.
  • save_to_json(): Stores the processed data in a human-readable, structured format that's easy to use for downstream tasks like data analysis or NLP modeling.

Customization Ideas

  • Swap out save_to_json for a CSV writer if you prefer tabular format.
  • Add error handling (try-except blocks) to gracefully handle malformed entries in the input file.
  • Integrate advanced NLP tools (like spaCy) for more sophisticated preprocessing (e.g., named entity recognition, dependency parsing).

内容的提问来源于stack exchange,提问作者Samuel Oyediran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:23:28