You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scanner读取文本文件并将Token传递至其他类进行计算?

How to Generate and Pass Tokens from Scanner to Other Classes for Text Stats

Hey there, let's break this down step by step—you're on the right track using Scanner to read the file, but we need to formalize what a "Token" is and how to pass those tokens along to your calculation classes.

First: Define Your Token Class

First, you need a clear Token class to represent the units you're extracting (like words, punctuation, or lines—depends on what stats you're calculating). Let's assume you're working with word tokens for this example:

public class Token {
    private String value;
    // Optional: Add a type enum if you need to distinguish words, punctuation, etc.
    // private TokenType type;

    public Token(String value) {
        this.value = value;
    }

    public String getValue() {
        return value;
    }

    // Optional: Override toString for easier debugging
    @Override
    public String toString() {
        return value;
    }
}

Modify Your Scanner Class to Generate Tokens

Your current analyzeBookText method is reading input, but instead of just storing strings, you'll create Token objects and send them to your calculation classes. You have two practical options here:

Option 1: Process tokens on the fly (no need to store all)

If you don't need to keep all tokens around for later use, pass each token to your stats class as you read it. This is memory-efficient for large files:

public class BookAnalyzer {
    private final StatsCalculator statsCalculator;

    // Inject your stats calculator (cleaner than instantiating inside the class)
    public BookAnalyzer(StatsCalculator statsCalculator) {
        this.statsCalculator = statsCalculator;
    }

    public void analyzeBookText(Scanner input) {
        // Adjust delimiter to split words and punctuation separately
        input.useDelimiter("\\s+|(?=[.,!?])|(?<=[.,!?])");
        
        while (input.hasNext()) {
            String tokenValue = input.next().trim();
            // Skip empty strings from trimming
            if (!tokenValue.isEmpty()) {
                Token token = new Token(tokenValue);
                statsCalculator.processToken(token);
            }
        }
        input.close();
    }
}

Option 2: Collect all tokens first (for multiple calculations)

If you need to reuse the tokens for different stats calculations, collect them into a list first:

public class BookAnalyzer {
    public List<Token> analyzeBookText(Scanner input) {
        List<Token> tokens = new ArrayList<>();
        input.useDelimiter("\\s+|(?=[.,!?])|(?<=[.,!?])");
        
        while (input.hasNext()) {
            String tokenValue = input.next().trim();
            if (!tokenValue.isEmpty()) {
                tokens.add(new Token(tokenValue));
            }
        }
        input.close();
        return tokens;
    }
}

Create a Stats Class to Process Tokens

Now build a class that takes tokens and computes your desired stats (word count, unique words, average word length, etc.):

public class StatsCalculator {
    private int totalWordCount = 0;
    private Set<String> uniqueWords = new HashSet<>();
    private int totalCharacterCount = 0;

    public void processToken(Token token) {
        String value = token.getValue();
        // Check if it's a word (skip punctuation/numbers if needed)
        if (value.matches("[a-zA-Z]+")) {
            totalWordCount++;
            uniqueWords.add(value.toLowerCase());
            totalCharacterCount += value.length();
        }
        // Add logic here for other token types if needed
    }

    // Getters to retrieve computed stats
    public int getTotalWordCount() {
        return totalWordCount;
    }

    public int getUniqueWordCount() {
        return uniqueWords.size();
    }

    public double getAverageWordLength() {
        return totalWordCount > 0 ? (double) totalCharacterCount / totalWordCount : 0;
    }
}

Putting It All Together

Here's how you'd use these classes in your main method:

public class Main {
    public static void main(String[] args) {
        // Option 1: Process tokens on the fly
        StatsCalculator stats = new StatsCalculator();
        BookAnalyzer analyzer = new BookAnalyzer(stats);
        
        try (Scanner scanner = new Scanner(new File("your-book.txt"))) {
            analyzer.analyzeBookText(scanner);
        } catch (FileNotFoundException e) {
            e.printStackTrace();
        }

        // Print your stats
        System.out.println("Total words: " + stats.getTotalWordCount());
        System.out.println("Unique words: " + stats.getUniqueWordCount());
        System.out.println("Average word length: " + stats.getAverageWordLength());

        // Option 2: Collect tokens first for multiple stats classes
        /*
        BookAnalyzer analyzer = new BookAnalyzer();
        List<Token> tokens = new ArrayList<>();
        try (Scanner scanner = new Scanner(new File("your-book.txt"))) {
            tokens = analyzer.analyzeBookText(scanner);
        } catch (FileNotFoundException e) {
            e.printStackTrace();
        }

        WordCountStats wordStats = new WordCountStats();
        wordStats.calculate(tokens);

        PunctuationStats punctuationStats = new PunctuationStats();
        punctuationStats.calculate(tokens);
        */
    }
}

Key Notes to Keep in Mind

  • Delimiter Tweaks: The regex delimiter I used splits on whitespace and separates punctuation from words (so "hello," becomes two tokens: "hello" and ","). Adjust this regex based on exactly what you consider a "token."
  • Resource Management: Always close your Scanner (or use try-with-resources like in the example to auto-close it and avoid resource leaks).
  • Token Types: If you need to distinguish between words, punctuation, numbers, etc., add an enum to the Token class and update your processing logic in the stats classes.

Content of the question originates from Stack Exchange, question author Alisia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:55:38