如何从文本文件中跳过无关内容,提取每行目标数值?
Hey Ben! Let's figure out how to pull all those numeric sequences (like your 123 or 000 placeholders) from each line of your large text file, ignoring all the other text. You started working with a Scanner—let's expand that approach and also cover another efficient method for bigger files.
Method 1: Using Scanner + Regular Expressions
This builds directly on the code you started writing. We'll read each line, then use a regex to find every numeric sequence in the line. This works great for most cases, and lets you choose whether to keep the numbers as strings (to preserve leading zeros like 000) or convert them to numeric types.
import java.io.File; import java.io.FileNotFoundException; import java.util.Scanner; import java.util.regex.Matcher; import java.util.regex.Pattern; public class NumberExtractor { public static void main(String[] args) { File inputFile = new File("your-file-path.txt"); // Regex to match one or more digits (covers 123, 000, etc.) Pattern numberPattern = Pattern.compile("\\d+"); try (Scanner scanner = new Scanner(inputFile)) { while (scanner.hasNextLine()) { String line = scanner.nextLine(); Matcher numberMatcher = numberPattern.matcher(line); // Loop through all numbers found in the current line while (numberMatcher.find()) { String numberString = numberMatcher.group(); // If you need an integer instead of a string: // int numericValue = Integer.parseInt(numberString); System.out.println(numberString); // Replace with your processing logic (e.g., save to a list) } } } catch (FileNotFoundException e) { System.err.println("Oops, couldn't find the file: " + e.getMessage()); e.printStackTrace(); } } }
Key Notes for This Method:
- The regex
\\d+matches any sequence of one or more digits—perfect for your123and000cases. - If you need to keep leading zeros (like keeping
000instead of converting it to0), stick with thenumberStringvariable. If you need a numeric type, useInteger.parseInt()orLong.parseLong()for larger numbers. - Adjust the regex if your numbers have decimals (use
\\d+\\.?\\d*) or negatives (use-?\\d+).
Method 2: Using BufferedReader for Large Files
If your text file is extremely large (hundreds of thousands of lines), BufferedReader can be more efficient than Scanner. The core logic stays the same—we just switch to a different file-reading approach:
import java.io.BufferedReader; import java.io.FileReader; import java.io.IOException; import java.util.regex.Matcher; import java.util.regex.Pattern; public class LargeFileNumberExtractor { public static void main(String[] args) { String filePath = "your-file-path.txt"; Pattern numberPattern = Pattern.compile("\\d+"); try (BufferedReader reader = new BufferedReader(new FileReader(filePath))) { String line; while ((line = reader.readLine()) != null) { Matcher numberMatcher = numberPattern.matcher(line); while (numberMatcher.find()) { System.out.println(numberMatcher.group()); // Add your custom processing here } } } catch (IOException e) { System.err.println("Error reading the file: " + e.getMessage()); e.printStackTrace(); } } }
Quick Tips:
- Don't forget to replace
your-file-path.txtwith the actual path to your text file. - If you need to collect all numbers instead of printing them, use a
List<String>orList<Integer>to store them as you extract them.
内容的提问来源于stack exchange,提问作者Ben

