Python文件读取:read()/readlines()方括号与圆括号访问差异及推荐
Hey there! Let's break down your questions about Python file operations, specifically the differences between square bracket indexing and calling read()/readlines() with parentheses, plus which approach to prefer.
Let’s break this down by each operation:
a. Square Bracket Indexing (e.g., lines[0])
Square brackets aren’t a file operation—they’re how you index into already loaded in-memory objects:
- When you run
lines = poem.read(),linesbecomes a single string containing the entire file content. Indexing it (likelines[0]) grabs individual characters from this string. - When you run
lines = poem.readlines(),linesbecomes a list where each element is a line from the file. Indexing it (likelines[0]) grabs an entire line from this list.
Key point: This indexing has nothing to do with the file’s "read pointer"—you’re just accessing data that’s already been pulled into memory.
b. Parentheses in read(n)
read(n) is a method directly on the file object, where n specifies the number of bytes to read:
- If
nis 0: Returns an empty string (no bytes are read). - If
nis a positive number: Reads up tonbytes starting from the file’s current read pointer. After each call, the pointer moves forward by the number of bytes read, so subsequent calls pick up where the last left off. - If you omit
n,read()reads the entire file at once (same asread(-1)).
Your example shows this perfectly:
poem.read(1) # 'A' (reads 1 byte, pointer moves to position 1) poem.read(2) # 'wa' (reads next 2 bytes, pointer moves to position 3) poem.read(3) # 'ke!' (reads next 3 bytes, pointer moves to position 6)
c. Parentheses in readlines(n)
This one is trickier—readlines(n) uses n as an approximate byte threshold:
- It reads full lines until the total number of bytes read exceeds
n. - If
nis 0: It reads all remaining lines in the file (since any content will exceed a 0-byte threshold). - The file pointer moves with each call, so subsequent calls read lines starting from where the last call ended.
Your examples illustrate this behavior:
- When you first open the file and call
readlines(1), the first line’s byte count is way over 1, so it returns just that line. - If you first call
readlines(0)(which reads all lines), the pointer hits the end of the file—soreadlines(1)returns an empty list (no lines left to read).
The core reason is what each operation targets:
- Square brackets operate on in-memory strings/lists: Results are consistent based on the loaded data, and don’t affect the file’s read position.
- Parentheses call file object methods: Results depend on the current read pointer position, and each call modifies that position for future operations.
It depends on your use case:
- For random access to specific characters/lines: Load the content into memory first with
read()orreadlines(), then use indexing. Note: Avoid this for very large files—it will consume too much RAM. - For processing large files line-by-line: Iterate directly over the file object (e.g.,
for line in poem:). This is memory-efficient because it reads one line at a time instead of loading the entire file. - For reading specific byte chunks: Use
read(n)—this is ideal for binary files or text files that need to be split into fixed-size segments. - For storing all lines in a list: Use
readlines()(but again, be cautious with large files).
内容的提问来源于stack exchange,提问作者user13714373

