問題文
A support team feeds product codes such as XR-4482B into a chatbot and notices that the usage report counts many more tokens than the number of words in their messages. A developer is asked to explain the counting. Which explanation is correct?
選択肢
- The report counts each character as a token, which is why a message containing long product codes produces a much larger number than a reader would expect.
- The report counts both the request and the response, so the number is always exactly twice the number of words the user wrote, and the product codes make no difference.
- The tokenizer splits rare or unusual strings into smaller pieces that exist in its vocabulary, so one written word can become several tokens.
- The tokenizer adds one token per punctuation mark and one per space, and the sum of those additions accounts for the entire difference seen in the report.