You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中自动将文本中的字符起止位置映射为对应单词位置?

Mapping Character Positions to Word Indices in Text

Got it, let's figure out how to convert character start/end positions into the corresponding word indices in a given text. Based on your example, we need to take a character range and find which words fall within (or overlap with) that range, then return their 0-based indices.

Step-by-Step Approach

  • Parse Input Data: Extract the text and convert the character position string into integer start/end values.
  • Track Word Boundaries: Split the text into words, and for each word, calculate its exact start and end character positions in the original text.
  • Match Character Range to Words: Identify which words' character ranges overlap with the input character position range, then collect their indices.

Python Implementation

Here's a working code example that handles your specific case and can be adapted to others:

# Input data from your example
input_data = {
    'text': 'Metiamide an histamine H2-receptors antagonist has been used to treat a case of Zollinger-Ellison syndrome characterized by a long standing diarrhea, an important gastric hypersecretion and a moderatly elevated plasma gastrin but without digestive ulceration.',
    'char_position': '[80.0, 106.0]'
}

# Step 1: Parse character positions to integers
start_char, end_char = map(int, input_data['char_position'].strip('[]').split(','))

# Step 2: Generate word list with their character boundaries
words_with_pos = []
current_char_pos = 0
for word in input_data['text'].split():
    word_start = current_char_pos
    word_end = current_char_pos + len(word)
    words_with_pos.append({
        'word': word,
        'start': word_start,
        'end': word_end
    })
    # Move past the word plus the following space
    current_char_pos = word_end + 1

# Step 3: Find overlapping word indices
word_indices = []
for idx, word_info in enumerate(words_with_pos):
    # Check if the word's range overlaps with the input character range
    if word_info['start'] <= end_char and word_info['end'] >= start_char:
        word_indices.append(idx)

# Build the output
output_data = {
    'text': input_data['text'],
    'word_position': f'[{", ".join(map(str, word_indices))}]'
}

print(output_data)
# Output: {'text': 'Metiamide an histamine H2-receptors antagonist has been used to treat a case of Zollinger-Ellison syndrome characterized by a long standing diarrhea, an important gastric hypersecretion and a moderatly elevated plasma gastrin but without digestive ulceration.', 'word_position': '[13, 14]'}

Key Notes

  • Word Splitting: We use split() which splits on whitespace, which works for your example (since target words are separated by spaces). If you need to handle punctuation attached to words (like "diarrhea,"), you might want to add preprocessing to clean up punctuation, but that depends on your specific use case.
  • Indexing: The example uses 0-based indices (which matches your desired output where Zollinger-Ellison is index 13 and syndrome is 14). If you need 1-based indices, just add 1 to each index in word_indices.

内容的提问来源于stack exchange,提问作者joasa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 22:27:32