如何在Python中自动将文本中的字符起止位置映射为对应单词位置?
Mapping Character Positions to Word Indices in Text
Got it, let's figure out how to convert character start/end positions into the corresponding word indices in a given text. Based on your example, we need to take a character range and find which words fall within (or overlap with) that range, then return their 0-based indices.
Step-by-Step Approach
- Parse Input Data: Extract the text and convert the character position string into integer start/end values.
- Track Word Boundaries: Split the text into words, and for each word, calculate its exact start and end character positions in the original text.
- Match Character Range to Words: Identify which words' character ranges overlap with the input character position range, then collect their indices.
Python Implementation
Here's a working code example that handles your specific case and can be adapted to others:
# Input data from your example input_data = { 'text': 'Metiamide an histamine H2-receptors antagonist has been used to treat a case of Zollinger-Ellison syndrome characterized by a long standing diarrhea, an important gastric hypersecretion and a moderatly elevated plasma gastrin but without digestive ulceration.', 'char_position': '[80.0, 106.0]' } # Step 1: Parse character positions to integers start_char, end_char = map(int, input_data['char_position'].strip('[]').split(',')) # Step 2: Generate word list with their character boundaries words_with_pos = [] current_char_pos = 0 for word in input_data['text'].split(): word_start = current_char_pos word_end = current_char_pos + len(word) words_with_pos.append({ 'word': word, 'start': word_start, 'end': word_end }) # Move past the word plus the following space current_char_pos = word_end + 1 # Step 3: Find overlapping word indices word_indices = [] for idx, word_info in enumerate(words_with_pos): # Check if the word's range overlaps with the input character range if word_info['start'] <= end_char and word_info['end'] >= start_char: word_indices.append(idx) # Build the output output_data = { 'text': input_data['text'], 'word_position': f'[{", ".join(map(str, word_indices))}]' } print(output_data) # Output: {'text': 'Metiamide an histamine H2-receptors antagonist has been used to treat a case of Zollinger-Ellison syndrome characterized by a long standing diarrhea, an important gastric hypersecretion and a moderatly elevated plasma gastrin but without digestive ulceration.', 'word_position': '[13, 14]'}
Key Notes
- Word Splitting: We use
split()which splits on whitespace, which works for your example (since target words are separated by spaces). If you need to handle punctuation attached to words (like "diarrhea,"), you might want to add preprocessing to clean up punctuation, but that depends on your specific use case. - Indexing: The example uses 0-based indices (which matches your desired output where Zollinger-Ellison is index 13 and syndrome is 14). If you need 1-based indices, just add 1 to each index in
word_indices.
内容的提问来源于stack exchange,提问作者joasa
相关产品推荐
相关产品推荐

