Python读取指定TXT文件提取特定数据并存入元组的需求
Extract Specific Bolded Sections from TXT File into Tuples with Python
Alright, let's figure out how to pull those highlighted c5t... entries from your text file and store them in tuples. The key here is targeting the pattern those entries follow—they all start with c5t, end with d0, and are followed by two capacity values (like 834G 118G).
Approach
We'll use regular expressions to identify and extract all matching entries, then format them into tuples (either as full string entries, or split into individual components if you need granular data).
Code Example
import re def extract_specific_data(file_path): # Regex pattern to match the bolded device + capacity entries pattern = r'c5t\w+d0 \d+(\.\d+)?[TG] \d+(\.\d+)?[TG]' with open(file_path, 'r') as file: content = file.read() # Grab all matching entries from the file content matches = re.findall(pattern, content) # Option 1: Store full entries as a tuple of strings full_entry_tuple = tuple(matches) # Option 2: Split each entry into (device_id, total_capacity, used_capacity) tuples split_component_tuple = tuple(tuple(entry.split()) for entry in matches) return full_entry_tuple, split_component_tuple # Usage example target_file = "your_target_file.txt" full_entries, split_entries = extract_specific_data(target_file) print("Full entries as a tuple:") print(full_entries) print("\nEntries split into component tuples:") print(split_entries)
Breakdown
- Regex Pattern: The pattern
r'c5t\w+d0 \d+(\.\d+)?[TG] \d+(\.\d+)?[TG]'is tailored to match:- The device identifier: Starts with
c5t, ends withd0, with alphanumeric characters in between - Capacity values: Supports whole numbers (like
834G) or decimals (like22.3G), ending with eitherT(terabyte) orG(gigabyte)
- The device identifier: Starts with
- File Reading: We open and read the file in one go for efficiency, then scan the entire content for matches.
- Tuple Conversion: Tuples are perfect here since your extracted data is fixed and immutable. We offer two formats: full string entries for simplicity, or split components if you need to access device IDs or capacities individually.
Sample Output
For your provided text snippet, the output would look like this:
Full entries as a tuple: ('c5t600507680C800000001CBd0 834G 118G', 'c5t600507680C00352d0 216G 22.3G', 'c5t600507680C800001CDd0 790G 162G', 'c5t600507680C8000000037Dd0 203G 34.8G') Entries split into component tuples: (('c5t600507680C800000001CBd0', '834G', '118G'), ('c5t600507680C00352d0', '216G', '22.3G'), ('c5t600507680C800001CDd0', '790G', '162G'), ('c5t600507680C8000000037Dd0', '203G', '34.8G'))
内容的提问来源于stack exchange,提问作者Mieszko Czerniak
相关产品推荐
相关产品推荐

