已编写脚本下载Wikipedia的pageviews.gz文件,解压后乱码如何读取?
Hey there! Let's work through how to read those Wikipedia pageviews files properly. The "unrecognizable content" you're seeing is almost always either a misstep in extraction, or not understanding the file's specific format structure. Here's a step-by-step breakdown to fix this:
Step 1: Confirm you've extracted the file correctly
First, double-check that you're actually working with the uncompressed text file, not still looking at the gzipped archive:
- Command line: Use
gzip -d pageviewsXXXXX.gz(orgunzip pageviewsXXXXX.gz) to decompress the file. This will replace the.gzfile with a plain text version. To verify it worked, runfile pageviewsXXXXX—you should see output likeASCII textorUTF-8 text. - GUI tools: If using 7-Zip, WinZip, or similar, make sure you extract the text file to your filesystem (not just preview it inside the archive). Preview panels sometimes display plain text incorrectly if they're optimized for binary files.
Step 2: Understand the pageviews file format
Once extracted, each line in the text file follows a consistent space-separated structure:
- Timestamp:
YYYYMMDDHH(e.g.,2024052013for May 20, 2024, 13:00 UTC) - Project code: Like
en.wikipediafor English Wikipedia,fr.wikipediafor French, etc. - Page title: URL-encoded (underscores replace spaces, special characters are escaped)
- View count: Integer of how many times the page was viewed that hour
- Total bytes transferred: Sum of bytes served for those views (optional in older files)
A sample line might look like this:2024052013 en.wikipedia Stack_Overflow 1234 567890
Step 3: Reading/parsing the file
Manual reading
Open the extracted text file with a plain text editor (Notepad++, VS Code, Sublime Text, or less/cat in the terminal). Avoid word processors like Microsoft Word—they'll try to apply formatting and mess up the display. For a quick peek in the terminal, run head -10 pageviewsXXXXX to see the first 10 lines.
Programmatic parsing
If you want to process the data with code, here are examples in common languages:
Python (read gzip directly, no extraction needed)
You can skip manual extraction entirely by using Python's built-in gzip module:
import gzip with gzip.open('pageviewsXXXXX.gz', 'rt', encoding='utf-8') as file: for line in file: # Split line into fields (handles one or more spaces as separators) fields = line.strip().split() if len(fields) >= 4: timestamp = fields[0] project = fields[1] page_title = fields[2].replace('_', ' ') # Convert underscores back to spaces view_count = int(fields[3]) bytes_served = int(fields[4]) if len(fields) == 5 else None # Example: Print data for English Wikipedia pages if project == 'en.wikipedia': print(f"Page: {page_title} | Views: {view_count}")
Bash (quick command-line analysis)
For fast filtering or aggregation without writing code:
- Count total entries:
gzip -dc pageviewsXXXXX.gz | wc -l - Filter English Wikipedia pages:
gzip -dc pageviewsXXXXX.gz | grep "^.*en.wikipedia.*" - Get top 10 most viewed pages:
gzip -dc pageviewsXXXXX.gz | sort -k4nr | head -10
Step 4: Troubleshooting garbled content
If the file still looks like binary garbage after proper extraction:
- Check download integrity: Verify the file's checksum (MD5/SHA) against the one provided where you downloaded the file. Mismatches mean the download was corrupted.
- Re-download: Partial or interrupted downloads often cause corruption—try getting the file again.
Alternative: Use the Wikimedia Pageviews API
If dealing with bulk files feels cumbersome, you can use the Wikimedia Pageviews API to fetch view data on demand. This lets you request hourly/daily/monthly views for specific pages, projects, or time ranges, and returns JSON responses that are easy to parse. No need to download or decompress large files!
For example, you can request data for a single page, filter by project, or get aggregated stats across multiple pages—all via simple HTTP requests.
内容的提问来源于stack exchange,提问作者Laerte Junior

